Live data from Hacker News

Why wordfreq will not be updated

github.com

311–320 of 542 posts

Re: Why wordfreq will not be updated

#311
post #145

Earlier quoted context omitted.

Perhaps they only find the increase shocking on the left. There's a bunch of ways to measure political opinions. The Authoritarian-liberal one being one of many. The economic-left and the economic-right are becoming more separated from the social-left and the social-right. Tribalism also causes people to take on the positions of their 'tribe' which may be distinct to what their own personalities might normally gravit…

Isn't that just what happens when you keep pushing the Overton Window to the right? What would have been 'centrists' have to become more authoritarian to stand ground or else they let their position get absorbed by the stronger leaning side. When one side refuses to compromise even slightly, you have two options: give in or dig in.

If the left-most edge of the Overton Window had been pushed right of the centrist position of, say, the 90s that would be correct. But it's obvious that isn't at all true -- many moderate left positions today would have been considered outlandishly radical in the 90s.

It's rather more likely that the increase of left-authoritarianism is due to increasing resistance to moving the Overton Window, both the left-most edge and the right-most edge, leftwards. As resistance increases more forceful techniques are necessary.

Re: Why wordfreq will not be updated

#312
post #265

Earlier quoted context omitted.

IMO HN actually scores quite highly in terms of health/politics and so forth content because the both mainstream and fringe ideas get both shown and pushback. A vaping discussion brought up glycerin used was safe and the same thing used in smoke machines and someone else brought up a study showing that smoke machines are an occasional safety issue. Nowhere near every discussion goes that well but stick around and you…

As a software engineer married to a healthcare professional, I disagree strongly about the quality of the healthcare discussions here. A whole lot of the conversation is software engineers who think that they can reason from first principles in two minutes about this thing that professionals dedicate their whole lives to mastering, and who therefore don't understand the most basic concepts of the field. Sometimes I t…

Spend time with medical researchers and they start disparaging Doctors. Everyone wants that one authoritative source free from bias, but IMO even having a few voices in the crowd worth listening to beats most other options.

Re: Why wordfreq will not be updated

#313
post #306
post #300

This is one of the vanguards warning of the changes coming in the post-AI world. >> Generative AI has polluted the data Just like low-background steel marks the break in history from before and after the nuclear age, these types of data mark the distinction from before and after AI. Future models will begin to continue to amplify certain statistical properties from their training, that amplified data will continue to…

> Future models will begin to continue to amplify certain statistical properties from their training, that amplified data will continue to pollute the public space from which future training data is drawn. That's why on FB I mark my own writing as AI generated, and the AI generated slop as genuine. Because what is disguised as "transparency disclaimer" is just flagging content of what's a potential dataset to train f…

I'm sorry for the low-content remark, but, oh my god... I never thought about doing this, and now my mind is reeling at the implications. The idea of shielding my own writing from AI-plagiarism by masquerading it as AI-generated slop in the first place... but then in the same stroke, further undermining our collective ability to identify genuine human writing, while also flagging my own work as low-value to my readers, hoping that they can read between the lines. It's a fascinating play.

Re: Why wordfreq will not be updated

#314

I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…

> If I can in any way purchase something without the web, I'mma do that

To get to the milk you'll have to walk by 3 rows of chips and soda.

Re: Why wordfreq will not be updated

#315
post #4

I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.

Congratulations on "shipping", I've had a background task to create pretty much exactly this site for a while. What is your cutoff date? I made this handy list, in research for mine:

  2017: Invention of transformer architecture
  June 2018: GPT-1
  February 2019: GPT-2
  June 2020: GPT-3
  March 2022: GPT-3.5
  November 2022: ChatGPT
You may want to add kiwix archives from before whatever date you choose. You can find them on the Internet Archive, and they're available for Wikipedia, Stack Overflow, Wikisource, Wikibooks, and various other wikis.

Re: Why wordfreq will not be updated

#316
post #314

I'm going to call it: The Web is dead. Thanks to "AI" I spend more time now digging through searches trying to find something useful than I did back in 2005. And the sites you do find are largely garbage. As a random example: just trying to find a particular popular set of wireless earbuds takes me at least 10 minutes, when I already know the company, the company's website, other vendors that sell the company's goods…

> If I can in any way purchase something without the web, I'mma do that To get to the milk you'll have to walk by 3 rows of chips and soda.

Yeah, this is why I still use the web to order things in a nutshell lol

Re: Why wordfreq will not be updated

#317

> Now the Web at large is full of slop generated by large language models, written by no one to communicate nothing. Fair and accurate. In the best cases the person running the model didn't write this stuff and word salad doesn't communicate whatever they meant to say. In many cases though, content is simply pumped out for SEO with no intention of being valuable to anyone.

[deleted]

Re: Why wordfreq will not be updated

#318
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

It's crazy to attribute the downfall of the web/search to Google. What does Google have to do with all the genuine open web content, Google's source of wealth, getting starved by (increasingly) walled gardens like Facebook, Reddit, Discord?

I don't see how Google's SEO rules being written or unwritten has any bearing. Spammers will always find a way.

Re: Why wordfreq will not be updated

#319
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

> I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. Blog spam was generally written by humans. While it sucked for other reasons, it seemed fine for measuring basic word frequencies in human-written tex…

Isn't it the other way around?

SEO text carefully tuned to tf-idf metrics and keyword stuffed to them empirically determined threshold Google just allows should have unnatural word frequencies.

LLM content should just enhance and cement the status quo word frequencies.

Outliers like the word "delve" could just be sentinels, carefully placed like trap streets on a map.

Re: Why wordfreq will not be updated

#320

Earlier quoted context omitted.

Isn't that just what happens when you keep pushing the Overton Window to the right? What would have been 'centrists' have to become more authoritarian to stand ground or else they let their position get absorbed by the stronger leaning side. When one side refuses to compromise even slightly, you have two options: give in or dig in.

If the left-most edge of the Overton Window had been pushed right of the centrist position of, say, the 90s that would be correct. But it's obvious that isn't at all true -- many moderate left positions today would have been considered outlandishly radical in the 90s. It's rather more likely that the increase of left-authoritarianism is due to increasing resistance to moving the Overton Window, both the left-most edg…

> many moderate left positions today would have been considered outlandishly radical in the 90s.

Such as?

Post reply on HN