Earlier quoted context omitted.
> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…
I don’t think they were talking about the quality of Google search results. I believe they were talking about how the data was processed by the wordfreq project.
Why wordfreq will not be updated
151–160 of 542 posts
Re: Why wordfreq will not be updated
#152> the Web at large is full of slop generated by large language models, written by no one to communicate nothing That’s neither fair nor accurate. That slop is ultimately generated by the humans who run those models; they are attempting (perhaps poorly) to communicate something . > two companies that I already despise Life’s too short to go through it hating others. > it's very likely because they are creating a plagi…
This kind of AI slop is quite literally written by no one (an algorithm pushed it out), and it doesn't communicate anything since communication first requires some level of understanding of the source material - and LLM's are just predicting the likely next token without understanding. I would also extend this to AI slop written by someone with a limited domain understanding, they themselves have nothing new to offer, nor the expertise or experience to ensure the AI is producing valuable content.
I would go even further and say it's "read by no one" - people are sick and tired of reading the next AI slop article on google and add stuff like "reddit" to the end of their queries to limit the amount of garbage they get.
Sure there are people using LLMs to enhance their research, but a vast, vast majority are using it to create slop that hits a word limit.
Re: Why wordfreq will not be updated
#153Earlier quoted context omitted.
Clever name. I like the analogy.
I don't seem to get it.
The analogy is that data is now contaminated with AI like steel is now contaminated with nuclear fallout.
https://en.wikipedia.org/wiki/Low-background_steel
>Low-background steel, also known as pre-war steel[1] and pre-atomic steel,[2] is any steel produced prior to the detonation of the first nuclear bombs in the 1940s and 1950s. Typically sourced from ships (either as part of regular scrapping or shipwrecks) and other steel artifacts of this era, it is often used for modern particle detectors because more modern steel is contaminated with traces of nuclear fallout.[3][4]
Re: Why wordfreq will not be updated
#154Re: Why wordfreq will not be updated
#155Earlier quoted context omitted.
Clever name. I like the analogy.
I don't seem to get it.
Re: Why wordfreq will not be updated
#156Earlier quoted context omitted.
Clever name. I like the analogy.
I don't seem to get it.
For applications that need to avoid the background radiation (like physics research), pre atomic age steel is extracted, like from old shipwrecks.
Re: Why wordfreq will not be updated
#157Re: Why wordfreq will not be updated
#158Re: Why wordfreq will not be updated
#159Earlier quoted context omitted.
And they're all flooded with low effort trash and useless. The only remaining reliable source - now that many newspapers are axing the remaining staff in favour of LLMs - is pre-2020 print cookbooks. Anything online or printed later must be assumed to be tainted, full of untested sewage and potentially dangerous suggestions.
Well there's https://www.allrecipes.com/author/chef-john/ on that particular site.
Re: Why wordfreq will not be updated
#160Earlier quoted context omitted.
This feels like a second, magnitudes larger Eternal September. I wonder how much more of this the Internet can take before everyone just abandons it entirely. My usage is notably lower than it was in even 2018, it's so goddamn hard to find anything worth reading anymore (which is why I spend so much damn time here, tbh).
I think it's an arms race, but it's an open question who wins. For a while I thought email as a medium was doomed, but spammers mostly lost that arms race. One interesting difference is that with spam, the large tech companies were basically all fighting against it. But here, many of the large tech companies are either providing tools to spammers (LLMs) or actively encouraging spammy behaviors (by integrating LLMs in…
There's probably an analogy to be made about the open decentralised internet in the age of AI here, if it gets to the point that search engines have to assume all sites are spam by default until proven otherwise, much like how an email server is assumed guilty until proven innocent.