Live data from Hacker News

Why wordfreq will not be updated

github.com

41–50 of 542 posts

Re: Why wordfreq will not be updated

#41
post #10

All those writers who'll soon be out of job and/or already are and basically unhireable for their previous tasks should be paid for by the AI hyperscalers to write anything at all on one condition: not a single sentence in their works should be created with AI. (I initially wanted to say 'paid for by the government' but that'd be socialising losses and we've had quite enough of that in the past.)

Who programs the tapes? https://en.wikipedia.org/wiki/Profession_(novella)

Re: Why wordfreq will not be updated

#42
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

It certainly feels like the amount of regurgitated, nonsensical, generated content (nontent?) has risen spectacularly specifically in the past few years. 2021 sounds about right based on just my own experience, even though I can't point to any objective source backing that up.

Re: Why wordfreq will not be updated

#43
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

Aunt may's brownie recipe (or at least her thoughts on it) are likely something you'd want if you want to reflect how humans use language. Both news-style and encyclopedia-style writing represent a pretty narrow slice.

Re: Why wordfreq will not be updated

#46
post #14

>"Now Twitter is gone anyway, its public APIs have shut down, and the site has been replaced with an oligarch's plaything, a spam-infested right-wing cesspool called X. Even if X made its raw data feed available (which it doesn't), there would be no valuable information to be found there. >Reddit also stopped providing public data archives, and now they sell their archives at a price that only OpenAI will pay. >And g…

> What beautiful doublethink. Given just how many AI bots scrape up everything they can, oftentimes ignoring robots.txt or any rate limits (there have been a few complaint threads on HN about that), I can hardly blame the operators of large online services just cutting off data feeds. Twitter however didn't stop their data feeds due to AI or because they wanted money, they stopped providing them because its new owner…

What was Reddit’s excuse? They did roughly the same thing (and have just as much garbage content).

In other words, why is it wrong for X but okay for Reddit? If you ignore one individual’s politics, the two services did the same thing.

Re: Why wordfreq will not be updated

#48

Man the AI folks really wrecked everything. Reminds me of when those scooter companies started just dumping their scooters everywhere without asking anybody if they wanted this.

perhaps germane to this thread, I think the scooter thing was an investment bubble. it was easier to burn investment money on new scooters than to collect and maintain old ones. until the money ran out.

Re: Why wordfreq will not be updated

#50

I wonder if anyone will fork the project. Apart from anything else, the data may still be useful given that we know it is polluted. In fact, it could act as a means of judging the impact of LLMs via that very pollution.

I guess it would be interesting but differentiating pollution from language evolution seems very tricky since getting a non polluted corpus gets harder and harder
Post reply on HN