Live data from Hacker News

Why wordfreq will not be updated

github.com

31–40 of 542 posts

Re: Why wordfreq will not be updated

#31
post #10

All those writers who'll soon be out of job and/or already are and basically unhireable for their previous tasks should be paid for by the AI hyperscalers to write anything at all on one condition: not a single sentence in their works should be created with AI. (I initially wanted to say 'paid for by the government' but that'd be socialising losses and we've had quite enough of that in the past.)

People have been paid to generate noise for a decade+ now. Garbage in, garbage out will always be true.

Next token-seeking is a solved problem. Novel thinking can be solved by humans and possibly by AI soon, but adding more garbage to the data won't improve things.

Re: Why wordfreq will not be updated

#32
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

Re: Why wordfreq will not be updated

#33
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Indexability is orthogonal to readability.

Re: Why wordfreq will not be updated

#36

It could be used to spot LLM generated text. compare the frequency of words to those used in human natural writings and you spot the computer from the human.

It could be used to differentiate LLM text from pre-LLM human text maybe. The thing, our AIs may not be very good at learning but our brains are. The more we use AI, the more we integrate LLMs and other tools into our life, the more their output will influence us. I believe there was a study (or a few anecdotes) where college papers checked for AI material were marked AI written even though they were written by humans because the students used AI during their studying and learned from it.

Re: Why wordfreq will not be updated

#37
post #22

Earlier quoted context omitted.

AI companies are indeed hiring such people to generate customized training data for them.

Is it the same companies that simply took all the writers' previous work (hoping to be billionaires before the courts understand)?

Yes. This was always the failure with the argument that copyright was the relevant issue... Once the model was proven out, we knew some wealthy companies would hire humans to generate the training data that the companies could then own in whole, at the relative expense of all other humans that didn't get paid to feed the machines.

Re: Why wordfreq will not be updated

#38

Enshittification is accelerating. A good 70% of my Facebook feed is now obviously AI generated images with AI generated text blurbs that have nothing to do with the accompanying images likely posted by overseas bot farms. I'm also noticing more and more "books" on Amazon that are clearly AI generated and self published.

It's okay. Amazon has limited authors to self publishing only 3 books per day (yes, really). That will surely solve the problem.

Re: Why wordfreq will not be updated

#39
"Multi-script languages

Two of the languages we support, Serbian and Chinese, are written in multiple scripts. To avoid spurious differences in word frequencies, we automatically transliterate the characters in these languages when looking up their words.

Serbian text written in Cyrillic letters is automatically converted to Latin letters, using standard Serbian transliteration, when the requested language is sr or sh."

I'd support keeping both scripts (српска ћирилица and latin script) , similarly to hiragana (ひらがな) and katakana (カタカナ) in Japanese.

Post reply on HN