Live data from Hacker News

Why wordfreq will not be updated

github.com

71–80 of 542 posts

Re: Why wordfreq will not be updated

#72
post #61
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

>And yet LLMs were still fed articles written for Googlebot, not humans. How do we know what content LLMs were fed? Isn't that a highly guarded secret? Won't the quality of the content be paramount to the quality of the generated output or does it not work that way?

We do know that the open web consitutes the bulk of the trainig data, although we don't get to know the specific webpages that got used. Plus some more selected sources, like books, of which again we only know that those are books but not which books were used. So it's just a matter of probability that there was a good amount of SEO spam as well.

Re: Why wordfreq will not be updated

#73
post #52

Earlier quoted context omitted.

That's why search engines rated them highly, and why a million spam sites cropped up that paid writers $1/essay to pretend to be Aunt May, and why today every recipe website has a gigantic useless fake essay in front of their copypasted made up recipes.

I hate how looking for recipes has become so… disheartening. Online recipes are fine for reputable sources like newspapers where professional recipe writers are paid for their contributions, but searching for some Aunt May's recipe for 'X' in the big ocean of the internet is pointless — too much raw sewage dumped in. It sucks, because sharing recipes seemed like one of those things the internet could be really good a…

There seem to be quite a few recipe sharing sites around - e.g. allrecipes.com.

Re: Why wordfreq will not be updated

#74
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

This feels like a second, magnitudes larger Eternal September. I wonder how much more of this the Internet can take before everyone just abandons it entirely. My usage is notably lower than it was in even 2018, it's so goddamn hard to find anything worth reading anymore (which is why I spend so much damn time here, tbh).

I think it's an arms race, but it's an open question who wins.

For a while I thought email as a medium was doomed, but spammers mostly lost that arms race. One interesting difference is that with spam, the large tech companies were basically all fighting against it. But here, many of the large tech companies are either providing tools to spammers (LLMs) or actively encouraging spammy behaviors (by integrating LLMs in ways that encourage people to send out text that they didn't write).

Re: Why wordfreq will not be updated

#75
post #46

Earlier quoted context omitted.

> What beautiful doublethink. Given just how many AI bots scrape up everything they can, oftentimes ignoring robots.txt or any rate limits (there have been a few complaint threads on HN about that), I can hardly blame the operators of large online services just cutting off data feeds. Twitter however didn't stop their data feeds due to AI or because they wanted money, they stopped providing them because its new owner…

What was Reddit’s excuse? They did roughly the same thing (and have just as much garbage content). In other words, why is it wrong for X but okay for Reddit? If you ignore one individual’s politics, the two services did the same thing.

Reddit shut their API access down only very recently, after the AI craze went off. Twitter did so right after Musk took over, way before Reddit, way before AI ever went nuts.

Re: Why wordfreq will not be updated

#76

Earlier quoted context omitted.

I hate how looking for recipes has become so… disheartening. Online recipes are fine for reputable sources like newspapers where professional recipe writers are paid for their contributions, but searching for some Aunt May's recipe for 'X' in the big ocean of the internet is pointless — too much raw sewage dumped in. It sucks, because sharing recipes seemed like one of those things the internet could be really good a…

There seem to be quite a few recipe sharing sites around - e.g. allrecipes.com.

And they're all flooded with low effort trash and useless.

The only remaining reliable source - now that many newspapers are axing the remaining staff in favour of LLMs - is pre-2020 print cookbooks. Anything online or printed later must be assumed to be tainted, full of untested sewage and potentially dangerous suggestions.

Re: Why wordfreq will not be updated

#77

I wonder if anyone will fork the project. Apart from anything else, the data may still be useful given that we know it is polluted. In fact, it could act as a means of judging the impact of LLMs via that very pollution.

I guess it would be interesting but differentiating pollution from language evolution seems very tricky since getting a non polluted corpus gets harder and harder

One way to tackle it would be to use LLMs to generate synthetic corpuses, so you have some good fingerprints for pollution. But even there I'm not sure how doable that is given the speed at which LLMs are being updated. Even if I know a particular page was created in, say, January 2023, I may no longer be able to try to generate something similar now to see how suspect it is, because the precise setups of the moment may no longer be available.

Re: Why wordfreq will not be updated

#78

I wonder if anyone will fork the project. Apart from anything else, the data may still be useful given that we know it is polluted. In fact, it could act as a means of judging the impact of LLMs via that very pollution.

I guess it would be interesting but differentiating pollution from language evolution seems very tricky since getting a non polluted corpus gets harder and harder

Arguably it is a form of language evolution. I bet humans have started using "delve" more too, on average. I think the best we can do is look at the trends and think about potential causes.

Re: Why wordfreq will not be updated

#79

Enshittification is accelerating. A good 70% of my Facebook feed is now obviously AI generated images with AI generated text blurbs that have nothing to do with the accompanying images likely posted by overseas bot farms. I'm also noticing more and more "books" on Amazon that are clearly AI generated and self published.

It's okay. Amazon has limited authors to self publishing only 3 books per day (yes, really). That will surely solve the problem.

Hah! I'm trying to figure out the exact date that crossed from "plausible line from a Stross or Sterling novel" [1] to "of course they did".

[1] Or maybe Sheckley or Lem, now that I think about it.

Re: Why wordfreq will not be updated

#80
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

> The training data is weighted by a quality metric

At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail and AI voice reading an AI script that was never graced by human eyes before being shat out onto the internet, and is of course always a glowing recommendation since the point is to get the viewer to click an affiliate link.

Google has been playing the SEO cat and mouse game forever, so can startups with a fraction of the experience be expected to do any better at filtering the noise out of fresh web scrapes?

Post reply on HN