Live data from Hacker News

LLMs can get "brain rot"

llm-brain-rot.github.io

201–210 of 310 posts

Re: LLMs can get "brain rot"

#201

Earlier quoted context omitted.

From the current WSJ front page: Paul Ingrassia's 'Nazi Streak' Musk Tosses Barbs at NASA Chie After SpaceX Criticism Travis Kelce Teams Up With Investor for Activist Campaign at Six Flags A Small North Carolina College Becomes a Magnet for Wealthy Students Cracker Barrel CEO Explains Short-Lived Logo Change If that's the benchmark for high quality training material we're in trouble.

In general I find WSJ articles very well written. It's not their fault if much of today's news is about clowns.

Their editorial department is an embarrassment imo. Sycophancy for conservatism thinly veiled as intellectualism.

Re: LLMs can get "brain rot"

#202

Earlier quoted context omitted.

I think this article has already made the rounds here, but I still think about it. I love using em dashes! It really makes me sad that I need to avoid them now to sound human https://bassi.li/articles/i-miss-using-em-dashes

Same here. I recently learned it was an LLM thing, and I've been using them forever. Also relevant: https://news.ycombinator.com/item?id=45226150

its not an llm thing -- its just -- folks don't know how to use them (pun intended).

Same for ; "" vs '', ex, eg, fe, etc. and so many more.

I like em all, but I'm crazy.

Re: LLMs can get "brain rot"

#203

Isn't this just garbage in garbage out with an attention grabbing title?

Yes, I am concerned about the Computer Science profession >"“Brain Rot” for LLMs isn’t just a catchy metaphor—it reframes data curation as cognitive hygiene for AI" A metaphor is exactly what it is because not only do LLMs not possess human cognition, there's certainly no established science of thinking they're literally valid subjects for clinical psychological assessment. How does this stuff get published, this is…

> How does this stuff get published

"published" only in the sense of "self-published on the Web". This manuscript has not or not yet been passed the peer review process, which is what scientist called "published" (properly).

Re: LLMs can get "brain rot"

#204
post #144

Earlier quoted context omitted.

Karpathy made a point recently that the random Common Crawl sample is complete junk, and that something like an WSJ article is extremely rare in it, and it's a miracle the models can learn anything at all.

>Turns out that LLMs learn a lot better and faster from educational content as well. This is partly because the average Common Crawl article (internet pages) is not of very high value and distracts the training, packing in too much irrelevant information. >The average webpage on the internet is so random and terrible it's not even clear how prior LLMs learn anything at all. You'd think it's random articles but it's n…

Don‘t forget the terabytes of torrented ebooks.

https://www.tomshardware.com/tech-industry/artificial-intell...

https://www.classaction.org/news/1.5b-anthropic-settlement-e...

Re: LLMs can get "brain rot"

#205

Earlier quoted context omitted.

> Those things are terrible; They look awful and destroy coherency of writing Totally agree. What the fuck did Nabokov, Joyce and Dickinson know about language. /s

Nothing. They wrote fiction.

I guess I'll ask: what's wrong with fiction?

Re: LLMs can get "brain rot"

#206

Earlier quoted context omitted.

I really could only be talking about LLMs but social media is also low quality. The quality (or lack of it) if such texts is self evident. If you are unable to discern that I can’t help you.

“The quality if such texts…” Indeed. The humans have bested the machines again.

I think that’s a good example of a superficial problem in a quickly typed statement, easily ignored, vs the profound and deep problems with LLM texts - they are devoid of meaning and purpose.

Re: LLMs can get "brain rot"

#207
post #47

“Studying “Brain Rot” for LLMs isn’t just a catchy metaphor—it reframes data curation as cognitive hygiene for AI, guiding how we source, filter, and maintain training corpora so deployed systems stay sharp, reliable, and aligned over time.” An LLM-written line if I’ve ever seen one. Looks like the authors have their own brainrot to contend with.

It is sad people study "brain rot" for LLMs but not for humans. If people were more engaged in cognitive hygiene for humans, many of the social media platforms would be very sane.

What do you base your claim on that people don't study that? I do not follow the research in that area but would find it highly unlikely there was no research into it.

Re: LLMs can get "brain rot"

#208

Earlier quoted context omitted.

I really could only be talking about LLMs but social media is also low quality. The quality (or lack of it) if such texts is self evident. If you are unable to discern that I can’t help you.

“The quality if such texts…” Indeed. The humans have bested the machines again.

Your comment was low quality noise while the one you replied to was on topic and useful. A short and useful comment with a typo is high quality content while a perfectly written LLM comment would be junk.

Re: LLMs can get "brain rot"

#210

Earlier quoted context omitted.

I think this article has already made the rounds here, but I still think about it. I love using em dashes! It really makes me sad that I need to avoid them now to sound human https://bassi.li/articles/i-miss-using-em-dashes

> I love using em dashes Keep using them. If someone is deducing from the use of an emdash that it's LLM produced, we've either lost the battle or they're an idiot. More pointedly, LLMs use emdashes in particular ways. Varying spacing around the em dash and using a double dash (--) could signal human writing.

The solution is clear: Unicode needs cryptographically signed dashes and whitespace characters.
Post reply on HN