Live data from Hacker News

A sleep-like consolidation mechanism for LLMs

arxiv.org

71–80 of 155 posts

Re: A sleep-like consolidation mechanism for LLMs

#71

To reach a more brain-like behavior LLMs need to integrate your inputs into their model dynamically, essentially retraining real-time based on the most salient input. Human brains do this selectively all the time and it's part of our plasticity. Biologically humans do similar compression, so introducing a similar concept to an LLM also feels reasonable. Hardware isn't fast/cheap enough to do this on an ongoing basis,…

>To reach a more brain-like behavior LLMs need to integrate your inputs into their model dynamically, essentially retraining real-time based on the most salient input.

That's already possible with LLMs. The challenge is that 1. it would allow permanently jail-breaking models and 2. there'd be no way for them to efficiently transfer what they'd learned to a new model generation.

Re: A sleep-like consolidation mechanism for LLMs

#72
post #20
post #14

Earlier quoted context omitted.

They provide an explanation for using the term "sleep": > In animals, the transfer from short-term memory to long-term memory is thought to be supported by hippocampal replay [33], especially during sleep [41]; in this phase, short-term hippocampal memories are reactivated and consolidated into cortical synaptic weights. Sleep makes animals unable to respond to external stimuli, suggesting that it must provide enough…

The function of sleep in animals is largely obscure. One thing we do know for certain is that it is necessary, it is needed in "dumb" animals as well as in you and I. If an animal can't sleep it will eventually die. I don't think that applies to the activity described in the OP. Does their LLM "die" if it can't perform the function described?

Is a volcano described as dormant (dormire, literally sleep) also inaccurate and deeply problematic? BTW, it's not anthropomorphized as sleep has existed long before humans.

"Sleep" is just used in their context to describe a non-interactive mode and they didn't lean heavily into zoomorphic - I think you mean - parallels.

You're grinding an axe on a single term. What is your broader hangup with them using the term "sleep"?

> Does their LLM "die" if it can't perform the function described?

We're reaching an age where LMGTFY should now be Let Me LLM That For You. Have you tried asking an LLM this question about the article? I believe it answers it very well.

Re: A sleep-like consolidation mechanism for LLMs

#73
post #4

I can't pretend to understand how LLMs work, but I can be sure that anthropomorphizing their functions is not helpful to an objective debate over their abilities. Does a motor vehicle get "sleep" when it is serviced? When I reboot a computer, is that equivalent to a nap?

I think it's interesting that folks are suddenly taking issue with "anthropomorphizing" language used in AI as if we haven't been doing this since the earliest days of computing (see "memory", "child", "parent", etc). It helps folks understand things at the correct level without needing domain knowledge

Re: A sleep-like consolidation mechanism for LLMs

#74
post #62

Earlier quoted context omitted.

Interesting that the scientific debate is settled, because you said so. Researchers who study prion diseases would probably be surprised to hear it.

Huh? Ask Claude or do some research on the topic if you don’t believe me. A prion disease killing you has nothing whatsoever to do with the lack of sleep. The insomnia is a side effect, not the cause. Jeez. People here are really stretching to defend their false “we die without sleep” claim.

Here's what Claude has to say about our exchange here.. since you asked.

> You're using absence of evidence as evidence of absence — which is a weak foundation when the evidence is genuinely hard to capture. You can't ethically deprive humans of sleep to death in a lab, and FFI affects only a handful of families worldwide.

> On the prion disease specifically: researchers haven't dismissed the role of sleep deprivation they've actively attempted to treat the insomnia in FFI patients on the hypothesis that it contributes to decline. That's not how a field behaves when it considers something a settled, irrelevant symptom.

> More broadly, "no human has ever died from lack of sleep" is an extraordinarily strong claim. To support it you'd need to rule out sleep deprivation as a factor in every candidate case and have a complete understanding of the mechanism. We have neither. The honest position is "we don't know" — not confident assertion in either direction.

Re: A sleep-like consolidation mechanism for LLMs

#75
post #4

I can't pretend to understand how LLMs work, but I can be sure that anthropomorphizing their functions is not helpful to an objective debate over their abilities. Does a motor vehicle get "sleep" when it is serviced? When I reboot a computer, is that equivalent to a nap?

[deleted]

Re: A sleep-like consolidation mechanism for LLMs

#76
post #20

Earlier quoted context omitted.

The function of sleep in animals is largely obscure. One thing we do know for certain is that it is necessary, it is needed in "dumb" animals as well as in you and I. If an animal can't sleep it will eventually die. I don't think that applies to the activity described in the OP. Does their LLM "die" if it can't perform the function described?

I don't think it's necessarily correct to think of sleep in terms of "it is necessary for animals or they will die". It might be more useful to think of it as "it was so useful that animals who slept outcompeted all the animals who didn't". Meaning: it might just provide a big advantage. I don't want to overextend and assume that any advantage extends to LLMs. That rest-and-recuperate advantage might also extend to L…

Sleep-like states exist in animals with nervous systems with a complexity above that found in flatworms, even snails sleep. Sleep therefore appears to be an essential characteristic of more complex biological nervous systems, i.e. biological computers, should you care to stretch the analogy. The more complex the nervous system, the greater the requirement for sleep.

What is described in the OP is therefore not a specific characteristic of sleep. It may however be a "useful" rhetorical device.

I do however object to the extensive use of such rhetorical tricks in the conversations that surround LLMs. For example, why does a consumer-grade LLM display "thinking" while it is actually sending data from my computer to some datacentre, processing it, and sending the result back? Equally, why does it output human-emotive phrases such as "sorry" when such computation is revealed to be incorrect?

Such rhetorical tricks, and more, likely underlie to a large degree the popularity of LLMs, despite their actual performance being clearly below what the rhetoric implies.

Re: A sleep-like consolidation mechanism for LLMs

#78
post #62

Earlier quoted context omitted.

Interesting that the scientific debate is settled, because you said so. Researchers who study prion diseases would probably be surprised to hear it.

Huh? Ask Claude or do some research on the topic if you don’t believe me. A prion disease killing you has nothing whatsoever to do with the lack of sleep. The insomnia is a side effect, not the cause. Jeez. People here are really stretching to defend their false “we die without sleep” claim.

Provide some evidence to back up you assertions. Don't tell someone else to do it for you.

Re: A sleep-like consolidation mechanism for LLMs

#79

Earlier quoted context omitted.

This is why I object to sleep() from unistd.h. What an anthropomorphizing notion. Didn't early unix programmers understand that a computer isn't a living creature and therefore isn't capable of sleep? They must have been really stupid!

Some of them were straight up psychopaths too, as evidenced by `kill()` !

Indeed and using SIGKILL is really cruel. At least with SIGTERM the process can say its goodbyes. /j

Re: A sleep-like consolidation mechanism for LLMs

#80
The "sleep" thing gives me the creeps so in my head I'm just going to think of it as the difference between "response time retrieval" and "background consolidation".

I do think it points at something bigger than just attention architecture: "memory" isn't just storage, and merely longer context isn't the same thing as having a better understanding of the source data.

I'm looking at this through the "personal AI" lens, where I think the missing "memory" layer seems to be consolidation & prioritization. It's not enough to just pattern match and grab the right emails, notes, etc, stuff them into the context window & hope, but instead it's useful to consider offline processing and turn events into durable state: clusters of observed data becomes episodes, assumptions, contradictions and power confidence for suggestions.

That also pushes up the need for provenance & inspectability. It's going to be interesting to see what kind of memory consolidation strategies are required for each domain use case.

Post reply on HN