Live data from Hacker News

Why wordfreq will not be updated

github.com

461–470 of 542 posts

Re: Why wordfreq will not be updated

#461
post #4

I created https://lowbackgroundsteel.ai/ in 2023 as a place to gather references to unpolluted datasets. I'll add wordfreq. Please submit stuff to the Tumblr.

:'( I thought I was clever for realising this parallel myself! Guess it's more obvious than I thought. Another example is how data on humans after 2020 or so can't be separated by sex because gender activists fought to stop recording sex in statistics on crime, medicine, etc.

I too realised this parallel and frequently tell people about it.

Edit: just the first one

Re: Why wordfreq will not be updated

#463

Earlier quoted context omitted.

It's worse because many humans don't know they are. I see a lot of outrage around fake posts already. People want to believe bad things from the other tribes. And we are going to feed them with it, endlessly.

Did you think the same thing when photoshop came out? It's relatively trivial to photoshop misinformation in a really powerful and undetectable way- but I don't see (legitimate) instances of groundbreaking news over a fake photo of the president or a CEO etc doing something nefarious. Why is AI different just because it's audio/video?

I did.

And it's not the grounbreaking the problem, it's the little constant lies.

Last week a photoshopped Musk tweet was going around, people getting all up in arms against it despite the fact it was very easy to spot as a fabricated one.

People didn't care, they hate the guy, they just wanted to fuel their hate more.

The whole planet run on fake content, magazin covers, food packaging, instagram pics of places that never looks that way...

And now, with AI, you can automate it and scale it up.

People are not ready. And in fact, they don't want to be.

Re: Why wordfreq will not be updated

#464
post #458

Earlier quoted context omitted.

On point 1, that’s surprising to me. A 2,000 word blog post would be 10 cents with GPT-4o. So you put out 1,000 of them, which is a lot, for $100.

But then you'll be competing for clicks with others who put out 1,000,000 posts for less costs because they used a small, self hosted model.

if you are a sales & marketing intern, have a potato laptop and $100 budget to spend on seo, you aren't going to be self hosting anything even if you know what that means.

Re: Why wordfreq will not be updated

#465
post #234

Earlier quoted context omitted.

I disagree. Even politics spurs intelligent, nuanced discussion here on HN. And to hold up discussions about MS as an example of 'extremely' low quality discussion is, ah, interesting. Do you have any recent examples of such discussions?

> spurs intelligent, nuanced discussion here on HN relative to what? reddit? also there's a trade off between entropy and "quality". too much "quality" and everyone gets bored and goes somewhere more entertaining

Relative to... unintelligent discussions?

I also don't care if people leave because HN isn't 'entertaining' enough. I don't come here for that, and I don't expect the community members that make this place what it is to either.

Re: Why wordfreq will not be updated

#466
I agree with the general ethos of the piece (albeit a few of the details are puzzling and unnecessarily partisan - content on X isn't invariably worthless drivel, nor does what Reddit is doing make much intellectual as opposed to economic [IPO-influenced] sense - but this line:

'OpenAI and Google can collect their own damn data. I hope they have to pay a very high price for it, and I hope they're constantly cursing the mess that they made themselves.'

really does betray some real naivete. OpenAI and Google could literally burn $10million dollars per day (okay, maybe not OpenAI - but Google surely could) and reasonably fail to notice. Whatever costs those companies have to pay to collect training data will be well worth it to them. Any messes made in the course of obtaining that data will be dealt with by an army of employees either manually cleaning up the data, or by algorithms Google has its own LLM write for itself.

I do find the general sense of impending dystopian inhumanity arising out of the explosion of LLMs to be super fascinating (and completely understandable).

Re: Why wordfreq will not be updated

#467
post #332
post #83

Earlier quoted context omitted.

I wish more people presented recipes like cooking for engineers. For example - Meat Lasagna https://www.cookingforengineers.com/recipe/36/Meat-Lasagna

I love the table-diagrams at the end. I've never seen anything like that until now and it really seems useful for visualization of the recipe and the sequence of steps.

Interestingly my wife has been writing recipes on post-it notes for years in that same style, with arrows instead of tables. And she's the opposite to an Engineer, a psychologist (interest in people vs objects).

When I saw them, they blew my mind. Short to store and easy to understand.

Re: Why wordfreq will not be updated

#468
> Now Twitter is gone anyway, its public APIs have shut down, and the site has been replaced with an oligarch's plaything, a spam-infested right-wing cesspool called X

God I hate this dystopic timeline we live in.

Re: Why wordfreq will not be updated

#469
post #383

Earlier quoted context omitted.

But you can already see it with Delve. Mistral uses "delve" more than baseline, because it was trained on GPT. So it's classic positive feedback. LLM uses delve more, delve appears in training data more, LLM uses delve more... Who knows what other semantic quirks are being amplified like this. It could be something much more subtle, like cadence or sentence structure. I already notice that GPT has a "tone" and Claude…

is the use of miscible here a clue? Or just some workplace vocabulary you've adapted analogically?

If you think that's niche wait til you hear about man-machine miscegenation

Re: Why wordfreq will not be updated

#470
post #460

Earlier quoted context omitted.

is the use of miscible here a clue? Or just some workplace vocabulary you've adapted analogically?

Human me just thought it was a good word for this. It implies some irreversible process of mixing, I think that characterizes this process really well.

There were dozens of 20th Century ideological movements which developed their own forms of "Newspeak" in their own native languages. Largely, natural human dialog between native speakers and between those opposed to the prevailing regime recoils violently at stilted, official, or just "uncool" usages in daily vernacular. So I wouldn't be too surprised to see a sharp downtick in the popular use of any word that becomes subject to an LLM's positive-feedback loop.

Far from saying the pool of language is now polluted, I think we now have a great data set to begin to discern authentic from inauthentic human language. Although sure, people on the fringes could get caught in a false positive for being bots, like you or I.

The biggest LLM of them all is the daily driver of all new linguistic innovation: Human society, in all its daily interactions. The quintillions of daily phrases exchanged and forever mutating around the globe - each mutation of phrase interacting with its interlocutor, and each drawing from not the last 500,000 tokens but the entire multi-modal, if you will, experience of each human to date in their entire lives - vastly eclipses anything any hardware could ever emulate given the current energy constraints. Software LLMs are just a state machine stuck in a moment in time. At best they will always lag, the way Stalinist language lagged years behind the patois of average Russians, who invented daily linguistic dodges to subvert and mock the regime. The same process takes place anywhere there is a dominant official or uncool accent or phrasing. The ghetto invents new words, new rhythm, and then it becomes cool in the middle class. The authorities never catch up, precisely because the use of subversive language is humanity's immune system against authority.

If there is one distinctly human trait, it's sniffing out anyone who sounds suspiciously inauthentic. (Sadly, it's also the trait that leads to every kind of conspiracy theorizing imaginable; but this too probably confers in some cases an evolutionary advantage). Sniffing out the sound of a few LLMs is already happening, and will accelerate geometrically, much faster than new models can be trained.

Post reply on HN