Live data from Hacker News

Why wordfreq will not be updated

github.com

161–170 of 542 posts

Re: Why wordfreq will not be updated

#162

Earlier quoted context omitted.

Clever name. I like the analogy.

I don't seem to get it.

From the blog

> Low Background Steel (and lead) is a type of metal uncontaminated by radioactive isotopes from nuclear testing. That steel and lead is usually recovered from ships that sunk before the Trinity Test in 1945.

Re: Why wordfreq will not be updated

#163
post #85

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

I never thought about that. Humans losing their ability to detect AI content from reality ? It's frightening.

Oh they definitely are. A lot of people are now calling out real photos as fake. I frequently get into stupid Instagram political arguments and a lot of times they come back with "yeah nice profile with all your AI art haha". It's all real high quality photography. Honestly, I don't think the avg person can tell anymore.

Re: Why wordfreq will not be updated

#164

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

> When I get asked if the person in a video is real, I still feel pretty confident to answer

I don't share your confidence in identifying real people anymore.

I often flag as "false-ish" a lot of things from genuinely real people, but who have adopted the behaviors of the TikTok/Insta/YouTube creator. Hell, my beard is grey and even I poked fun at "YouTube Thumbnail Face" back in 2020 in a video talk I gave. AI twigs into these "semi-human" behavioral patterns super fast and super hard.

There is a video floating around with pairs of young ladies with "This is real"/"This is not real" on signs. They could be completely lying about both, and I really can't tell the difference. All of them have behavioral patterns that seems a little "off" but are consistent with the small number of "influencer" videos I have exposure to.

Re: Why wordfreq will not be updated

#165

> Now the Web at large is full of slop generated by large language models, written by no one to communicate nothing. Fair and accurate. In the best cases the person running the model didn't write this stuff and word salad doesn't communicate whatever they meant to say. In many cases though, content is simply pumped out for SEO with no intention of being valuable to anyone.

[flagged]

The problem is that for the vast majority of use, LLM output is not revised or edited, and very many times I'm convinced the output wasn't even fully read.

Re: Why wordfreq will not be updated

#166
post #80

Earlier quoted context omitted.

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…

> Google has been playing the SEO cat and mouse game forever, so can startups with a fraction of the experience be expected to do any better at filtering the noise out of fresh web scrapes?

Google has been _monetizing_ the SEO game forever. They chose not to act against many notorious actors because the metric they optimize for is ad revenue and and those sites were loaded with ads. As long as advertisers didn’t stop buying, they didn’t feel much pressure to make big changes.

A smaller company without that inherent conflict of interest in its business model can do better because they work on a fundamentally different problem.

Re: Why wordfreq will not be updated

#167

Earlier quoted context omitted.

[flagged]

The problem is that for the vast majority of use, LLM output is not revised or edited, and very many times I'm convinced the output wasn't even fully read.

I assume FrustratedMonky's comment was satirical, given that it appears to have been written like an LLM and starts with a "but, but, but" which is how you might represent someone you disagree with presenting their argument.

Re: Why wordfreq will not be updated

#168

[flagged]

> Also, it is shocking how authoritarian the “left” has become in my lifetime. We are going through a general uptick in authoritarian "discussions" online. It's interesting that you are only seeing it on the "left".

They said nothing about it being “only” on the left.

I somewhat expect authoritarianism on the right and therefore would hold the left (to which I belong) at a higher standard.

Re: Why wordfreq will not be updated

#169

It could be used to spot LLM generated text. compare the frequency of words to those used in human natural writings and you spot the computer from the human.

it may work for a short time, but after a while natural language will evolve due to natural exposure of those new words or word patterns and even human will write in ways that, while being different from the LLMs, will also be different from the snapshot captured by this snapshot. It's already the case that we used to write differently 20 years ago from 50 years ago and even more so 100 years ago, etc

Re: Why wordfreq will not be updated

#170
post #103

Earlier quoted context omitted.

‘delve’ is given as an example right there in TFA.

Yes, but the material presented in no way makes distiction between potential organic growth of 'delve' vs. LLM induced use. They just note that even though 'delve' was on the rise, in 23-24 the word gains more popularity, at the same time ChatGPT rose. Word adoption is certainly not a linear phenomenon. And as the author states 'I don't think anyone has reliable information about post-2021 language usage by humans' S…

Even granting that we can disregard a really huge factor here, which I'm not sure we really can, one can not know beforehand how the clustering of the vocabulary is going to go pre-training, and its speculated that both at the center and at the edges of clusters we get random particularities. Hence the "solidgoldmagikarp" phenomenon and many others.
Post reply on HN