Live data from Hacker News

Why wordfreq will not be updated

github.com

101–110 of 542 posts

Re: Why wordfreq will not be updated

#101
post #68

I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated? AI assisted sure, but is that a problem if a human is in the mix, curating? I certainly have not encountered enough straight drivel where I would think it would have a significant effect on overall word statistics. I suspect there may be some over-identification of AI content happenin…

> I hear this complaint often but in reality I have encountered fairly little content in my day to day that has felt fully AI generated?

How confident are you in this assessment?

> straight drivel

We're past the point where what AI generates is "straight drivel"; every minute, it's harder to distinguish AI output from actual output unless you're approaching expertise in the subject being written about.

> a team of copywriters who pumped out mountains the most inane keyword packed text designed for literally no one but Google to read.

And now a machine can generate the same amount of output in 30 seconds. Scale matters.

Re: Why wordfreq will not be updated

#102
post #85

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

I never thought about that. Humans losing their ability to detect AI content from reality ? It's frightening.

It's worse because many humans don't know they are.

I see a lot of outrage around fake posts already. People want to believe bad things from the other tribes.

And we are going to feed them with it, endlessly.

Re: Why wordfreq will not be updated

#104

[flagged]

It did feel emotive but this wasn't the main point. Data is harder to get (or more expansive) and more polluted.

Felt super emotive to me, the problems the author is outlining, a) might not be an actual problems b) just require new thinking to solve

Re: Why wordfreq will not be updated

#105

[flagged]

Even if this is true it being an emotional decision, so much of twitter/X itself is now AI slop anyway, so it'd be worth it to just not even include it whether it was right wing or not.

Regardless, the owner is well within his right to make an emotional decision based on his beliefs to stop anyway.

Re: Why wordfreq will not be updated

#106

Intuitively I feel like word frequency would be one of the things least impacted by LLM output, no?

It'd be in fact quite the opposite. There comes a turning point where the majority of language usage would actually be written by AI, at which point we'd no longer be analysing the word frequency/usage by actual humans and so it wouldn't be representative of how humans actually communicate.

Or potentially even more dystopian would be that AI slop would be dictating/driving human communication going forward.

Re: Why wordfreq will not be updated

#108
post #74

Earlier quoted context omitted.

This feels like a second, magnitudes larger Eternal September. I wonder how much more of this the Internet can take before everyone just abandons it entirely. My usage is notably lower than it was in even 2018, it's so goddamn hard to find anything worth reading anymore (which is why I spend so much damn time here, tbh).

I think it's an arms race, but it's an open question who wins. For a while I thought email as a medium was doomed, but spammers mostly lost that arms race. One interesting difference is that with spam, the large tech companies were basically all fighting against it. But here, many of the large tech companies are either providing tools to spammers (LLMs) or actively encouraging spammy behaviors (by integrating LLMs in…

Another problem with this arms race is that spam emails actually are largely separable from ham emails for most people... or at least they were, for most of their run. The thousandth email that claims the UN has set aside money for me due to my non-existent African noble ancestry that they can't find anyone to give it to and I just need to send the Thailand embassy some money to start processing my multi-million yuan payout and send it to my choice of proxy in Colombia to pick it up is quite different from technical conversation about some GitHub issue I'm subscribed to, on all sorts of metrics.

However, the frontline of the email war has shifted lately. Now the most important part of the war is being fought over emails that look just like ham, but aren't. Business frauds where someone convinces you that they are the CEO or CFO or some VP and they need you to urgently buy this or that for them right now no time to talk is big business right now, and before you get too high-and-mighty about how immune you are to that, they are now extremely good at looking official. This war has not been won yet, and to a large degree, isn't something you necessarily win by AI either.

I think there's an analogy here to the war on content slop. Since what the content slop wants is just for you to see it so they can serve you ads, it doesn't need anything else that our algorithms could trip on, like links to malware or calls to action to be defrauded, or anything else. It looks just like the real stuff, and telling that it isn't could require a human rather vast amounts of input just to be mostly sure. Except we don't have the ability to authenticate where it came from. (There is no content authentication solution that will work at scale. No matter how you try to get humans to "sign their work" people will always work out how to automate it and then it's done.) So the one good and solid signal that helps in email is gone for general web content.

I don't judge this as a winning scenario for the defenders here. It's not a total victory for the attackers either, but I'd hesitate to even call an advantage for one side or the other. Fighting AI slop is not going to be easy.

Re: Why wordfreq will not be updated

#109
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Indexability is orthogonal to readability.

It should be, but sadly it’s not.
Post reply on HN