All those writers who'll soon be out of job and/or already are and basically unhireable for their previous tasks should be paid for by the AI hyperscalers to write anything at all on one condition: not a single sentence in their works should be created with AI. (I initially wanted to say 'paid for by the government' but that'd be socialising losses and we've had quite enough of that in the past.)
Why wordfreq will not be updated
41–50 of 542 posts
Re: Why wordfreq will not be updated
#42I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…
Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.
Re: Why wordfreq will not be updated
#43I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…
Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.
Re: Why wordfreq will not be updated
#44Re: Why wordfreq will not be updated
#45Re: Why wordfreq will not be updated
#46>"Now Twitter is gone anyway, its public APIs have shut down, and the site has been replaced with an oligarch's plaything, a spam-infested right-wing cesspool called X. Even if X made its raw data feed available (which it doesn't), there would be no valuable information to be found there. >Reddit also stopped providing public data archives, and now they sell their archives at a price that only OpenAI will pay. >And g…
> What beautiful doublethink. Given just how many AI bots scrape up everything they can, oftentimes ignoring robots.txt or any rate limits (there have been a few complaint threads on HN about that), I can hardly blame the operators of large online services just cutting off data feeds. Twitter however didn't stop their data feeds due to AI or because they wanted money, they stopped providing them because its new owner…
In other words, why is it wrong for X but okay for Reddit? If you ignore one individual’s politics, the two services did the same thing.
Re: Why wordfreq will not be updated
#47[flagged]
You’re not edgy dude. Give it up.
Re: Why wordfreq will not be updated
#48Man the AI folks really wrecked everything. Reminds me of when those scooter companies started just dumping their scooters everywhere without asking anybody if they wanted this.
Re: Why wordfreq will not be updated
#49Re: Why wordfreq will not be updated
#50I wonder if anyone will fork the project. Apart from anything else, the data may still be useful given that we know it is polluted. In fact, it could act as a means of judging the impact of LLMs via that very pollution.