Live data from Hacker News

A sleep-like consolidation mechanism for LLMs

arxiv.org

131–140 of 155 posts

Re: A sleep-like consolidation mechanism for LLMs

#131
post #4

I can't pretend to understand how LLMs work, but I can be sure that anthropomorphizing their functions is not helpful to an objective debate over their abilities. Does a motor vehicle get "sleep" when it is serviced? When I reboot a computer, is that equivalent to a nap?

First, this is not a "debate over the abilities" of LLMs. It's a proposed method to improve their performance, and the authors are free to call it however they think it makes sense. Second, explicitly avoiding things that sound like anthropomorphisation is equally not helpful- why avoid a metaphor that works? Third, it's really a pity that this pointless nitpicking is dominating the thread.

>this pointless nitpicking is dominating the thread.

What is dominating the thread are claims that the LLM operation in question is analogous to the function of sleep in humans. It obviously is not.

The anthropomorphization of LLMs has reached ridiculous proportions. Applying the same standards as used in this field to others would result in claims that laundry machines "hallucinated" that they had had sufficient water when they failed due to the faucet being turned off.

Re: A sleep-like consolidation mechanism for LLMs

#133
post #120

The entire industry is so desperate to anthropomorphize. What the paper describes is an offline recurrent consolidation phase: the model runs multiple forward passes over recently accumulated context, updates persistent fast weights in SSM blocks, then clears the KV cache before continuing. It has absolutely nothing to do with sleeping, but I believe the authors had a goal in mind when creating this title, and it was…

It is a descriptive analogy, get over yourself.

An intelligent reply from an obviously intelligent guy!

A more appropriate title would have been something like "Offline Recurrent Memory Consolidation for Long-Context Language Models". This is supposed to be a research paper, not a story book. The title should give context to other researchers, and not be clearly engineered for clicks. If you don't think so, that's your prerogative, but you're objectively wrong.

Re: A sleep-like consolidation mechanism for LLMs

#134

Earlier quoted context omitted.

It’s a bunch of Claude blather, and I love Claude. Just not worth copying over to HN, because the rush to get to a narrow answer to a narrow question elides the meaningful bits, ex. what does happen during sleep deprivation. Has a “not even wrong” air simply because you’re trying to get to true/false on a narrow question then pushing your research assistant to disavow what you’re quote unquote “skeptical” of.

This is little more than a fancy way of saying “Nu uh.” Such arguments are hardly convincing.

I don't understand this post. I can read it, but I don't understand it.

Why do you think I'm arguing something? Again, smacks of Pauli's not even wrong. You're confusing your headspace with everyone else's and rushed to copy pasta an AI you're browbeating into disavowing things you're skeptical of to...win an argument, I guess? Based on this post? Unclear to me what the argument is, or, if I'm understanding correctly, due to the narrow focus while being self-absorbed.

Re: A sleep-like consolidation mechanism for LLMs

#135

The idea of periodically stopping to write blocks of recent context into a fast-weight state is interesting, but I think it liked it better when E2E-TTT[1] did it. It's a more flexible and elegant continuous learning approach. Essentially it goes "You know how your model can remember its training data? Well, what if you treated its recent context like more training data and updated (some of) the weights using (mostly…

I wonder if we can get children to make something their life’s dream if we make the cool books about it when they are growing up? I wonder how flexible the human mind can be in convincing itself that it is fulfilling its dream?

Re: A sleep-like consolidation mechanism for LLMs

#137

The idea of periodically stopping to write blocks of recent context into a fast-weight state is interesting, but I think it liked it better when E2E-TTT[1] did it. It's a more flexible and elegant continuous learning approach. Essentially it goes "You know how your model can remember its training data? Well, what if you treated its recent context like more training data and updated (some of) the weights using (mostly…

Each model needs to be a separate copy, or at least have those particular weights be interchangeable, for every single user.

Remember Microsoft Tay.

https://en.wikipedia.org/wiki/Tay_(chatbot)#Initial_release

Re: A sleep-like consolidation mechanism for LLMs

#138
post #120

Earlier quoted context omitted.

It is a descriptive analogy, get over yourself.

An intelligent reply from an obviously intelligent guy! A more appropriate title would have been something like "Offline Recurrent Memory Consolidation for Long-Context Language Models". This is supposed to be a research paper, not a story book. The title should give context to other researchers, and not be clearly engineered for clicks. If you don't think so, that's your prerogative, but you're objectively wrong.

You write the paper, you write the title. So much anger over a title, you are graydon, make this about yourself.

Re: A sleep-like consolidation mechanism for LLMs

#139
post #90

Earlier quoted context omitted.

"Despite myriad studies, there is still no consensus on why sleep is needed for survival." https://www.nature.com/articles/d41586-025-00964-w (2025)

Probably because it evolved very early (like before bilateral symmetry, multi layer body cavity, or kidneys early... maybe even before multicellular animals early) and so has been incorporated as an essential pillar into multiple processes layered on top of that fundamental architecture.

You've merely stated observations about the context and the process that led to it. That doesn't in any way answer the question of what it's actually doing that's so essential.

Re: A sleep-like consolidation mechanism for LLMs

#140

The idea of periodically stopping to write blocks of recent context into a fast-weight state is interesting, but I think it liked it better when E2E-TTT[1] did it. It's a more flexible and elegant continuous learning approach. Essentially it goes "You know how your model can remember its training data? Well, what if you treated its recent context like more training data and updated (some of) the weights using (mostly…

I wonder if we can get children to make something their life’s dream if we make the cool books about it when they are growing up? I wonder how flexible the human mind can be in convincing itself that it is fulfilling its dream?

This sounds like a horror novel
Post reply on HN