Live data from Hacker News

Why wordfreq will not be updated

github.com

351–360 of 542 posts

Re: Why wordfreq will not be updated

#351
I understand the frustration shared in this post but I wholeheartedly disagree with the overall sentiment that comes with it.

The web isn't dead, (Gen)AI, SEO, spam and pollution didn't kill anything.

The world is chaotic and net entropy (degree of disorder) of any isolated or closed system will always increase. Same goes for the web. We just have to embrace it and overcome the challenges that come with it.

Re: Why wordfreq will not be updated

#352
post #349

Earlier quoted context omitted.

Oh they definitely are. A lot of people are now calling out real photos as fake. I frequently get into stupid Instagram political arguments and a lot of times they come back with "yeah nice profile with all your AI art haha". It's all real high quality photography. Honestly, I don't think the avg person can tell anymore.

I've reached a point where even if my first reaction to a photo is to be impressed, I then quickly think "oh but what it this is AI?" and then immediately my excitement for the photo is ruined because it may not actually be a photo at all.

I don't get that perspective at all. Who cares what made it.

Re: Why wordfreq will not be updated

#353
post #351

I understand the frustration shared in this post but I wholeheartedly disagree with the overall sentiment that comes with it. The web isn't dead, (Gen)AI, SEO, spam and pollution didn't kill anything. The world is chaotic and net entropy (degree of disorder) of any isolated or closed system will always increase. Same goes for the web. We just have to embrace it and overcome the challenges that come with it.

Here is an expert saying there is a problem and how it killed its research effort, and yet you say that things are the same as ever and nothing was killed.

Re: Why wordfreq will not be updated

#354
post #13

I agree in general but the web was already polluted by Google's unwritten SEO rules. Single-sentence paragraphs, multiple keyword repetitions and focus on "indexability" instead of readability, made the web a less than ideal source for such analysis long before LLMs. It also made the web a less than ideal source for training. And yet LLMs were still fed articles written for Googlebot, not humans. ML/LLM is the second…

Prior to Google we had Altavista and in those days it was incredibly common to find keywords spammed hundreds of times in white text on a white background in the footer of a page. SEO spam is not new, it's just different.

Re: Why wordfreq will not be updated

#355

Earlier quoted context omitted.

There are a series of challenges like: https://www.nytimes.com/interactive/2024/09/09/technology/ai... https://www.nytimes.com/interactive/2024/01/19/technology/ar... These are a little bit unfair, in that we're comparing handpicked examples, but I don't think many experts will pass a test like this. Technology only moves forward (and seemingly, at an accelerating pace). What's a little shocking to me is the speed of…

+100w chargers are one of the products I prefer to spend a little more on, so I get something from a company that knows it can be sued if they make a product that burns down your house or fries your phone. Flashlights? Sure, bring on aliexpress. USB cables with pop-off magnetically attached heads, no problem. But power supplies? Welp, to each their own!

And then you plug your cheap pop-off USB cable into the expensive 100w charger?

Re: Why wordfreq will not be updated

#356
post #55

Earlier quoted context omitted.

It certainly feels like the amount of regurgitated, nonsensical, generated content (nontent?) has risen spectacularly specifically in the past few years. 2021 sounds about right based on just my own experience, even though I can't point to any objective source backing that up.

SEO grifters have fully integrated AI at this point, there are dozens of turn-key "solutions" for mass-producing "content" with the absolute minimum effort possible. It's been refined to the point that scraping material from other sites, running it through the LLM blender to make it look original, and publishing it on a platform like Wordpress is fully automated end-to-end.

Or check out "money printer" on github: a tongue in cheek mashup of various tools to take a keyword as input and produce a youtube video with subtitles and narration as output.

Re: Why wordfreq will not be updated

#357
post #343

Earlier quoted context omitted.

As a software engineer married to a healthcare professional, I disagree strongly about the quality of the healthcare discussions here. A whole lot of the conversation is software engineers who think that they can reason from first principles in two minutes about this thing that professionals dedicate their whole lives to mastering, and who therefore don't understand the most basic concepts of the field. Sometimes I t…

> about this thing that professionals dedicate their whole lives to mastering After doing some healthcare work I ended up understanding that some topics are not well known even by the professionals dedicating their whole lives to that because there are big gaps in the human knowledge on the topics. I agree that people that think they can reason in two minutes about anything are a problem, but it's not a healthcare on…

As to the uncertainty and mysteries, you are 100% correct. One of the big failure modes for engineers in dealing with human health is the assumption that things are as simple and logical as the stuff we build, when it's simply not at all like that. There are (1) big arguments over basic things like "why do SSRI's work?" Outside of LLM's I can't think of a thing in software where we are still arguing about why things work in production. We never say "Why does Postgres work?" in the same way. (2)

And yes, this is true for many other areas of discussion at HN. It's just that it is most obvious to me in the area that my wife specializes in, because I pick up enough via osmosis from her to know when other people don't even have my limited level of understanding.

1: Or at least were 15 years ago when my wife told me about it- the argument might have been largely concluded and she just never updated me since I don't keep up with the medical literature the way she does.

2: Two decades ago there was a huge push for the "human genome project" under the basis that this would be "reading the blueprints for human life" and that would give us all of these medical breakthroughs. Basically none of those breakthroughs happened because we've spent the past 20 years learning all of the different ways that it is NOT a blueprint and that cells do things very differently from human engineers.

Re: Why wordfreq will not be updated

#358
A few years ago I began an effort to write a new tech book. I planned orig to do as much of it as I could across a series of commits in a public GitHub repo of mine.

I then changed course. Why? I had read increasing reports of human e-book pirates (copying your book's content then repackaging it for sale under a diff title, byline, cover, and possibly at a much lower or even much higher price.)

And then the rise of LLMs and their ravenous training ingest bots -- plagiarism at scale and potentially even easier to disguise.

"Not gonna happen." - Bush Sr., via Dana Carvey

Now I keep the bulk of my book material non-public during dev. I'm sure I'll share a chapter candidate or so at some point before final release, for feedback and publicity. But the bulk will debut all together at once, and only once polished and behind a paywall

Re: Why wordfreq will not be updated

#359
post #4

I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.

That's exactly the opposite of what the author wanted IMO. The author no more wants to be a part of this mess. Aggregating these sources would just makes it so much more easier for the tech giants to scrape more data.

The main concerns expressed in Robyn's note, as I read them, seem to be 1) generative AI has polluted the web with text that was not written by humans, and so it is no longer feasible to produce reliable word frequency data that reflects how humans use natural language; and 2) simultaneously, sources of natural language text that were previously accessible to researchers are now less accessible because the owners of that content don't want it used by others to create AI models without their permission. A third concern seems to be that support for and practice of any other NLP approaches is vanishing.

Making resources like wordfreq more visible won't exacerbate any of these concerns.

Re: Why wordfreq will not be updated

#360
post #351

I understand the frustration shared in this post but I wholeheartedly disagree with the overall sentiment that comes with it. The web isn't dead, (Gen)AI, SEO, spam and pollution didn't kill anything. The world is chaotic and net entropy (degree of disorder) of any isolated or closed system will always increase. Same goes for the web. We just have to embrace it and overcome the challenges that come with it.

I'm not so optimistic. The most basic requirements are:

1. Prove the human-ness of an author... 2. ...without grossly encroaching on their privacy. 3. Ensure that the author isn't passing off AI-generated material as their own.

We'll leave out the "don't let AI models train on my data" part for now.

Whatever solution we come up with, if any, will necessarily be mired in the politics of privacy, anonymity, and/or DRM. In any case, it's hard to conceive of a world where the human web returns as we once knew it.

Post reply on HN