Live data from Hacker News

LLMs can get "brain rot"

llm-brain-rot.github.io

161–170 of 310 posts

Re: LLMs can get "brain rot"

#161

Earlier quoted context omitted.

Suddenly I see all these people come out of the woodworks talking about "em dashes". Those things are terrible; They look awful and destroy coherency of writing. No wonder LLM's use them.

> Those things are terrible; They look awful and destroy coherency of writing Totally agree. What the fuck did Nabokov, Joyce and Dickinson know about language. /s

Nothing. They wrote fiction.

Re: LLMs can get "brain rot"

#162

Earlier quoted context omitted.

If they wanted to watermark (I always felt it is irresponsible not to, if someone wants to circumvent it that's on them) - they could use strategically placed whitespace characters like zero-width spaces, maybe spelling something out in Morse code the way genius.com did to catch google crawling lyric (I believe in that case it was left and right handed aposterofes)

Which could be removed with a simple filter. em dashes require at least a little bit of code to replace with their correct grammar equivalents.

> em dashes require at least a little bit of code to replace with their correct grammar equivalents

Or an LLM that could run on Windows 98. The em dashes--like AI's other annoyingly-repetitive turns of phrase--are more likely an artefact.

Re: LLMs can get "brain rot"

#164

Earlier quoted context omitted.

> Those things are terrible; They look awful and destroy coherency of writing Totally agree. What the fuck did Nabokov, Joyce and Dickinson know about language. /s

Nothing. They wrote fiction.

> Nothing

/s?

> They wrote fiction

Now do Carl Sagan and Richard Feynman.

Re: LLMs can get "brain rot"

#165

Earlier quoted context omitted.

Don't forget the "it's not just X, it's Y" formulation and the rule of 3.

More signs of AI Writing: https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing

Can we back this into the internet communities or corpuses of human work that excessively used these phrases? The "it's not just X" seems copy pasted from SEO marketing copy. But some of the others are less obvious.

Re: LLMs can get "brain rot"

#166

Earlier quoted context omitted.

> It’s a text generator regurgitating plausible phrases without understanding and producing stale and meaningless pablum. Are you describing LLM's or social media users? Dont conflate how the content was created with its quality. The "You must be at least this smart (tall) to publish (ride)" sign got torn down years ago. Speakers corner is now an (inter)national stage and it written so it must be true...

I really could only be talking about LLMs but social media is also low quality. The quality (or lack of it) if such texts is self evident. If you are unable to discern that I can’t help you.

“The quality if such texts…”

Indeed. The humans have bested the machines again.

Re: LLMs can get "brain rot"

#167

Earlier quoted context omitted.

Karpathy made a point recently that the random Common Crawl sample is complete junk, and that something like an WSJ article is extremely rare in it, and it's a miracle the models can learn anything at all.

From the current WSJ front page: Paul Ingrassia's 'Nazi Streak' Musk Tosses Barbs at NASA Chie After SpaceX Criticism Travis Kelce Teams Up With Investor for Activist Campaign at Six Flags A Small North Carolina College Becomes a Magnet for Wealthy Students Cracker Barrel CEO Explains Short-Lived Logo Change If that's the benchmark for high quality training material we're in trouble.

There is very, very little written work that will stand the test of time. Maybe the real bitter lesson is that training data quality is inversely proportional to scale and the technical capabilities exist but can never be realized

Re: LLMs can get "brain rot"

#168

Earlier quoted context omitted.

I think this article has already made the rounds here, but I still think about it. I love using em dashes! It really makes me sad that I need to avoid them now to sound human https://bassi.li/articles/i-miss-using-em-dashes

Same here. I recently learned it was an LLM thing, and I've been using them forever. Also relevant: https://news.ycombinator.com/item?id=45226150

> I’ve been using them forever.

Many other HN contributors have, too. Here’s the pre-ChatGPT em dash leaderboard:

https://www.gally.net/miscellaneous/hn-em-dash-user-leaderbo...

Re: LLMs can get "brain rot"

#169

Earlier quoted context omitted.

That is indeed an LLM-written sentence — not only does it employ an em dash, but also lists objects in a series — twice within the same sentence — typical LLM behavior that renders its output conspicuous, obvious, and readily apparent to HN readers.

hehe, I see what you did there.

it is amusing to use AI to write that...

Re: LLMs can get "brain rot"

#170

Earlier quoted context omitted.

That is indeed an LLM-written sentence — not only does it employ an em dash, but also lists objects in a series — twice within the same sentence — typical LLM behavior that renders its output conspicuous, obvious, and readily apparent to HN readers.

I think this article has already made the rounds here, but I still think about it. I love using em dashes! It really makes me sad that I need to avoid them now to sound human https://bassi.li/articles/i-miss-using-em-dashes

I still use them all the time, and if someone objects to my writing over them then I've successfully avoided having to engage with a dweeb.

(But in practice, I don't think I've had a single person suggest that my writing is LLM-generated despite the presence of em-dashes, so maybe the problem isn't that bad.)

Post reply on HN