Live data from Hacker News

Why wordfreq will not be updated

github.com

151–160 of 542 posts

Re: Why wordfreq will not be updated

#151
post #140
post #80

Earlier quoted context omitted.

> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…

I don’t think they were talking about the quality of Google search results. I believe they were talking about how the data was processed by the wordfreq project.

I was actually referring to the data ingestion for training LLMs, I don't know what filtering or weighting might be done with wordfreq.

Re: Why wordfreq will not be updated

#152

> the Web at large is full of slop generated by large language models, written by no one to communicate nothing That’s neither fair nor accurate. That slop is ultimately generated by the humans who run those models; they are attempting (perhaps poorly) to communicate something . > two companies that I already despise Life’s too short to go through it hating others. > it's very likely because they are creating a plagi…

This is just the "guns don't shoot people, people do." argument except in this case we quite literally have a massive upside incentive to remove people from the process entirely (i.e. websites that automatically generate new content every day) - so I don't buy it.

This kind of AI slop is quite literally written by no one (an algorithm pushed it out), and it doesn't communicate anything since communication first requires some level of understanding of the source material - and LLM's are just predicting the likely next token without understanding. I would also extend this to AI slop written by someone with a limited domain understanding, they themselves have nothing new to offer, nor the expertise or experience to ensure the AI is producing valuable content.

I would go even further and say it's "read by no one" - people are sick and tired of reading the next AI slop article on google and add stuff like "reddit" to the end of their queries to limit the amount of garbage they get.

Sure there are people using LLMs to enhance their research, but a vast, vast majority are using it to create slop that hits a word limit.

Re: Why wordfreq will not be updated

#153

Earlier quoted context omitted.

Clever name. I like the analogy.

I don't seem to get it.

Steel without nuclear contamination is sought after, and only available from pre-war / pre-atomic sources.

The analogy is that data is now contaminated with AI like steel is now contaminated with nuclear fallout.

https://en.wikipedia.org/wiki/Low-background_steel

>Low-background steel, also known as pre-war steel[1] and pre-atomic steel,[2] is any steel produced prior to the detonation of the first nuclear bombs in the 1940s and 1950s. Typically sourced from ships (either as part of regular scrapping or shipwrecks) and other steel artifacts of this era, it is often used for modern particle detectors because more modern steel is contaminated with traces of nuclear fallout.[3][4]

Re: Why wordfreq will not be updated

#155

Earlier quoted context omitted.

Clever name. I like the analogy.

I don't seem to get it.

It's a reference to the practise of scavenging steel from sources that were produced before nuclear testing began, as any steel produced afterwards is contaminated with nuclear isotopes from the fallout. Mostly ship wrecks, and WW2 means there are plenty of those. The pun in question is that his project tries to source text that hasn't been contaminated with AI generated material.

https://en.m.wikipedia.org/wiki/Low-background_steel

Re: Why wordfreq will not be updated

#156

Earlier quoted context omitted.

Clever name. I like the analogy.

I don't seem to get it.

After the detonation of the first nuclear weapons, any newly produced steel has a low dose of nuclear fallout.

For applications that need to avoid the background radiation (like physics research), pre atomic age steel is extracted, like from old shipwrecks.

https://en.m.wikipedia.org/wiki/Low-background_steel

Re: Why wordfreq will not be updated

#157

Earlier quoted context omitted.

Their batteries on the other hand…

Sure, they're worse than walking or biking, but compared to an electric car battery or an ICE car?

At least where I'm from, scooters have mostly replaced walking and biking, not car trips :(

Re: Why wordfreq will not be updated

#159
post #76

Earlier quoted context omitted.

And they're all flooded with low effort trash and useless. The only remaining reliable source - now that many newspapers are axing the remaining staff in favour of LLMs - is pre-2020 print cookbooks. Anything online or printed later must be assumed to be tainted, full of untested sewage and potentially dangerous suggestions.

Well there's https://www.allrecipes.com/author/chef-john/ on that particular site.

Chef John is the best.

Re: Why wordfreq will not be updated

#160
post #74

Earlier quoted context omitted.

This feels like a second, magnitudes larger Eternal September. I wonder how much more of this the Internet can take before everyone just abandons it entirely. My usage is notably lower than it was in even 2018, it's so goddamn hard to find anything worth reading anymore (which is why I spend so much damn time here, tbh).

I think it's an arms race, but it's an open question who wins. For a while I thought email as a medium was doomed, but spammers mostly lost that arms race. One interesting difference is that with spam, the large tech companies were basically all fighting against it. But here, many of the large tech companies are either providing tools to spammers (LLMs) or actively encouraging spammy behaviors (by integrating LLMs in…

The fight against spam email also led to mass consolidation of what was supposed to be a decentralised system though. Monoliths like Google and Microsoft now act as de-facto gatekeepers who decide whether or not you're allowed to send emails, and there's little to no transparency or recourse to their decisions.

There's probably an analogy to be made about the open decentralised internet in the age of AI here, if it gets to the point that search engines have to assume all sites are spam by default until proven otherwise, much like how an email server is assumed guilty until proven innocent.

Post reply on HN