Live data from Hacker News

Why wordfreq will not be updated

github.com

171–180 of 542 posts

Re: Why wordfreq will not be updated

#171
Ok so post author is AI skeptic and this is his retaliation, likely because his work is affected. I believe governments should address the problem with welfare but being against technical advances is always being in the wrong side of history.

Re: Why wordfreq will not be updated

#172
post #136

Earlier quoted context omitted.

Generative AI is inherently a political issue, its not surprising at all. There is the case of what is "truth". As soon as you start to ensure some quality of truth to what is generated, that is political. As soon as generative AI has the capability to take someone's job, that is political. The instant AI can make someone money, it is political. When AI is trained on something that someone has created, and now they c…

Then .. everything is political?

It is. Unfortunately.

Re: Why wordfreq will not be updated

#173
post #65

"Multi-script languages Two of the languages we support, Serbian and Chinese, are written in multiple scripts. To avoid spurious differences in word frequencies, we automatically transliterate the characters in these languages when looking up their words. Serbian text written in Cyrillic letters is automatically converted to Latin letters, using standard Serbian transliteration, when the requested language is sr or s…

Why is this a HN comment on a thread about it ending due to AI pollution?

[flagged]

Re: Why wordfreq will not be updated

#174

Earlier quoted context omitted.

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

It certainly feels like the amount of regurgitated, nonsensical, generated content (nontent?) has risen spectacularly specifically in the past few years. 2021 sounds about right based on just my own experience, even though I can't point to any objective source backing that up.

Upvoted for "nontent" alone: it'll be my go-to term from now on, and I hope it catches on.

Is it of your own coinage? When the AI sifts through the digital wreckage of the brief human empire, they may give you the credit.

Re: Why wordfreq will not be updated

#175

I guess a manageable, still-useful alternative would be to curate a whitelist of sources that don't use AI, and without making that list public, derive the word frequencies from only those sources. How to compile that list is left as an exercise for the reader. The result would not be as accurate as a broad sample of the web, but in a world where it's impossible to trust a broad sample of the web, it the option you a…

> curate a whitelist of sources that don't use AI,

I like this.

Maybe even take it a step further - have a badge on the source that is both human and machine visible to indicate that the content is not AI generated.

Re: Why wordfreq will not be updated

#176
post #76

Earlier quoted context omitted.

And they're all flooded with low effort trash and useless. The only remaining reliable source - now that many newspapers are axing the remaining staff in favour of LLMs - is pre-2020 print cookbooks. Anything online or printed later must be assumed to be tainted, full of untested sewage and potentially dangerous suggestions.

Well there's https://www.allrecipes.com/author/chef-john/ on that particular site.

I absolutely love Chef John. Great recipes and the cadence of his speech on YouTube (foodwishes) is very soothing, while he cooks up something amazing. If you're a home cook I highly recommend his recipes and his channel.

Re: Why wordfreq will not be updated

#177

"I don't think anyone has reliable information about post-2021 language usage by humans." We've been past the tipping point when it comes to text for some time, but for video I feel we are living through the watershed moment right now. Especially smaller children don't have a good intuition on what is real and what is not. When I get asked if the person in a video is real, I still feel pretty confident to answer but…

There are a series of challenges like: https://www.nytimes.com/interactive/2024/09/09/technology/ai... https://www.nytimes.com/interactive/2024/01/19/technology/ar... These are a little bit unfair, in that we're comparing handpicked examples, but I don't think many experts will pass a test like this. Technology only moves forward (and seemingly, at an accelerating pace). What's a little shocking to me is the speed of…

+100w chargers are one of the products I prefer to spend a little more on, so I get something from a company that knows it can be sued if they make a product that burns down your house or fries your phone.

Flashlights? Sure, bring on aliexpress. USB cables with pop-off magnetically attached heads, no problem. But power supplies? Welp, to each their own!

Re: Why wordfreq will not be updated

#178
post #41
post #10

All those writers who'll soon be out of job and/or already are and basically unhireable for their previous tasks should be paid for by the AI hyperscalers to write anything at all on one condition: not a single sentence in their works should be created with AI. (I initially wanted to say 'paid for by the government' but that'd be socialising losses and we've had quite enough of that in the past.)

Who programs the tapes? https://en.wikipedia.org/wiki/Profession_(novella)

_Thank you_. I read this story probably around 1980 (I think in a magazine that was subsequently trashed or garage-saled), and I have spent my adult life remembering the bones of the story, but not the author or the title.

Re: Why wordfreq will not be updated

#179
post #103

Earlier quoted context omitted.

‘delve’ is given as an example right there in TFA.

Yes, but the material presented in no way makes distiction between potential organic growth of 'delve' vs. LLM induced use. They just note that even though 'delve' was on the rise, in 23-24 the word gains more popularity, at the same time ChatGPT rose. Word adoption is certainly not a linear phenomenon. And as the author states 'I don't think anyone has reliable information about post-2021 language usage by humans' S…

The fact that making this distinction is impossible is reason enough to stop.

Re: Why wordfreq will not be updated

#180
post #80

Earlier quoted context omitted.

Yes but not quite as far as you imply. The training data is weighted by a quality metric, articles written by journalists and wikipedia contributors are given more weight than Aunt May's brownie recipe and corpoblogspam.

> The training data is weighted by a quality metric At least in Googles case, they're having so much difficulty keeping AI slop out of their search results that I don't have much faith in their ability to give it an appropriately low training weight. They're not even filtering the comically low-hanging fruit like those YouTube channels which post a new "product review" every 10 minutes, with an AI generated thumbnail…

Google has those problems because the company's revenue source (Ads) and the thing that puts it on the map (Search) are fundamentally at odds with one another.

A useful Search would ideally send a user to the site with the most signal and the fewest noise. Meanwhile, ads are inherently noise; they're extra pieces of information inserted into a webpage that at best tangentially correlate to the subject of a page.

Up until ~5 years ago, Google was able to strike a balance on keeping these two stable; you'd get results with some Ads but the signal generally outweighed the noise. Unfortunately from what I can tell from anecdotes and courtroom documents, the Ad team at Google has essentially hijacked every other aspect of the company by threatening that yearly bonuses won't be given out if they don't kowtow to the Ad teams wishes to optimize ad revenue somewhere in 2018-2019 and has no sign of stopping since there's no effective competition to Google. (There's like, Bing and Kagi? Nobody uses Bing though and Kagi is only used by tech enthusiasts. The problem with Google is that to copy it, you need a ton of computing resources upfront and are going up against a company with infinitely more money and ability to ensure users don't leave their ecosystem; go ahead and abandon Search, but good luck convincing others to give up say, their Gmail account, which keeps them locked to Google and Search will be there, enticing the average user.)

Google has absolutely zero incentive to filter out generative AI junk from their search results outside the amount of it that's damaging their PR since most of the SEO spam is also running Google Ads (since unless you're hosting adult content, Google's ad network is practically the only option). Their solution therefore isn't to remove the AI junk, but to instead reduce it enough to the degree where a user will not get the same type of AI junk twice.

Post reply on HN