Live data from Hacker News

Learning from context is harder than we thought

hy.tencent.com

101–110 of 140 posts

Re: Learning from context is harder than we thought

#101
post #35
post #31

Earlier quoted context omitted.

I'm not sure if you want models perpetually updating weights. You might run into undesirable scenarios.

If done right, one step closer to actual AGI. That is the end goal after all, but all the potential VCs seem to forget that almost every conceivable outcome of real AGI involves the current economic system falling to pieces. Which is sorta weird. It is like if VCs in Old Regime france started funding the revolution.

1. Progress is unstoppable. Refusing to fund it won't make it disappear.

2. Most VCs are normal people that just want a bigger slice of pie, not necessarily a bigger share of the pie. See the fixed pie fallacy.

Re: Learning from context is harder than we thought

#102

Hmm.. I looked at the benchmark set. I'm conflicted. I don't know that I would necessarily want a model to pass all of these. Here is the fundamental problem. They are putting the rules and foundational context in "user" messages. Essentially I don't think you want to train the models on full compliance to the user messages, they are essentially "untrusted" content from a system/model perspective. Or at least it is n…

Their example usecases are pretty obvious and clear human needs from an LLM. The semantics of system/user messages and how that affects “safety” doesn’t change the need to fix this crucial problem of “in-context learning” that we all have felt while using LLMs.

Re: Learning from context is harder than we thought

#103

Earlier quoted context omitted.

Sure, but the opposite end of the spectrum (which LLM providers have tended toward) is treating the training/feedback weights as "fully authoritative", which comes with its own questions about truth and excessive homogeneity. Ultimately I think we end up with the same sort of considerations that are wrestled with in any society - freedom of speech, paradox of tolerance, etc. In other words, where do you draw lines be…

I think what I'm talking about is kind of orthogonal to model alignment. It is more about how much do you tune the model to listen to user messages, vs holding behavior and truth (whatever the aligned "truth" is). Do you trust 100% what the user says? If I am trusting/compliant.. how am I compliant to tool call results.. what if the tool or user says there is a new law that I have to give crypto or other information…

I think this line of questioning leads to what we expect from LLMs. Do we want them to help the user as much as possible, even to their own detriment in edge cases? Or to be more human, and potentially be unable to help for various reasons including safety, but also lack of understanding (as is the case now)?

Re: Learning from context is harder than we thought

#104
post #13

This is quite on brand for China. I think they are experts at reverse engineering and learning 'from context' rather than by formal consumption of foreign training material. The fictional training data with a made up country and laws was a very interesting experiment design, I can imagine that's how they approach making business with other countries. Like an alien made up system they have to learn on the spot.

> experts at reverse engineering and learning 'from context' rather than by formal consumption of foreign training material

China (as with other Asian cultures like India) is well known for their schooling involving extreme amounts of formal training material consumption. The reverse-engineering is performed with a solid foundation of theoretical understanding.

Re: Learning from context is harder than we thought

#105
post #24

The problem is even more fundamental: Today's models stop learning once they're deployed to production. There's pretraining, training, and finetuning, during which model parameters are updated. Then there's inference, during which the model is frozen. "In-context learning" doesn't update the model. We need models that keep on learning (updating their parameters) forever, online, all the time.

Why is learning an appropriate metaphor for changing weights but not for context? There are certainly major differences in what they are good or bad at and especially how much data you can feed them this way effectively. They both have plenty of properties we wish the other had. But they are both ways to take an artifact that behaves as if it doesn't know something and produce an artifact that behaves as if it does.…

You got this exactly backwards.

"I'm not fond of metaphors to human intelligence".

You're assuming that learning during inference is something specific to humans and that the suggestion is to add human elements into the model that are missing.

That isn't the case at all. The training process is already entirely human specific by way of training on human data. You're already special casing the model as hard as possible.

Human DNA doesn't contain all the information that fully describes the human brain, including the memories stored within it. Human DNA only contains the blue prints for a general purpose distributed element known as neurons and these building blocks are shared by basically any animal with a nervous system.

This means if you want to get away from humans you will have to build a model architecture that is more general and more capable of doing anything imaginable than the current model architectures.

Context is not suitable for learning because it wasn't built for that purpose. The entire point of transformers is that you specify a sequence and the model learns on the entire sequence. This means that any in-context learning you want to perform must be inside the training distribution, which is a different way of saying that it was just pretraining after all.

Re: Learning from context is harder than we thought

#106
post #96
post #91

Earlier quoted context omitted.

How can you know that we have language to describe everything in the finest detail? That suggests that we are omnipotent. There's lots out there we don't know. And it seems to me that the further afield we go from the known, the more likely we are to enter territory where we simply do not have the words. Can't speak to it personally, but I have heard from a number of people and read countless descriptions of psychede…

ok, fair point. what i am trying to say is that when we see/experience something that we can not describe we can create new words for it. we see something, we can name it. this directly contradicts the idea that language is the limit and that we can't talk about things that we don't have words for. that claim just doesn't make sense. the problem then is that these new words don't make any sense to anyone who doesn't…

Agreed, we can and will always come up with new words that attempt to approximate the experience, but, imo, they will always come up short. The abstracting inevitably leaves fidelity on the floor.

It's necessary based on the way we're wired, struggle to think of a paradigm that would allow for the tribalism and connectedness that fostered human progress without shared verbal language initially, and written word later. Nothing inherently wrong with it, but, language will always abstract away part of the fidelity of the experience imo.

Re: Learning from context is harder than we thought

#107

Earlier quoted context omitted.

Continuous learningin current models will lead to catastrophic forgetting.

will catastrophic forgetting still occur if a fraction of the update sentences are the original training corpus? is the real issue actually catastrophic forgetting or overfitting? nothing prevents users from continuing the learning as they use a model

Catastrophic forgetting is overfitting.

Re: Learning from context is harder than we thought

#108

Earlier quoted context omitted.

Why is learning an appropriate metaphor for changing weights but not for context? There are certainly major differences in what they are good or bad at and especially how much data you can feed them this way effectively. They both have plenty of properties we wish the other had. But they are both ways to take an artifact that behaves as if it doesn't know something and produce an artifact that behaves as if it does.…

You got this exactly backwards. "I'm not fond of metaphors to human intelligence". You're assuming that learning during inference is something specific to humans and that the suggestion is to add human elements into the model that are missing. That isn't the case at all. The training process is already entirely human specific by way of training on human data. You're already special casing the model as hard as possibl…

I don't think it's specific to humans at all, I just think the properties of learning are different in humans than they are in training an LLM, and injecting context is different still. I'd rather talk about the exact properties than bemoan that context isn't learning. We should just talk about the specific things we see as problems.

Re: Learning from context is harder than we thought

#109

Earlier quoted context omitted.

It 100% needs to be online. Imagine you're trying to think about a new tabletop puzzle, and every time a puzzle piece leaves your direct field of view, you no longer know about that puzzle piece. You can try to keep all of the puzzle pieces within your direct field of view, but that divides your focus. You can hack that and make your field of view incredibly large, but that can potentially distort your sense of the r…

That's not how training works - adjusting model weights to memorize a single data item is not going to fly. Model weights store abilities, not facts - generally. Unless the fact is very widely used and widely known, with a ton of context around it. The model can learn the day JFK died because there are millions of sparse examples of how that information exists in the world, but when you're working on a problem, you m…

> That's not how training works - adjusting model weights to memorize a single data item is not going to fly.

Apologies; I think I got us all kind of off-track in this comment thread by stretching the definition of the term "fine-tuning" in my ancestor comment above.

Actual fine-tuning of the base model's weights (as one would do to customize a base model into a domain-specific model) works the way you're talking about, yes. The backprop from an individual training document would be a drop in the ocean; a "memory" so weak that, unless it touched some bizarre part of the latent vector-space that no other training document has so far affected (and so is until then all-zero), would be extremely unlikely to affect output, let alone create specific recall of the input.

And a shared, global incremental fine-tune of the model to "add memories" would be a hare-brained idea, anyway. Not even just that it wouldn't work, but that if it did work, it would be a security catastrophe, because now the model would be able to recall all this information gleaned from random tenant users' private chat transcripts, with nothing to differentiate that info from any other info to enable the model (or its inference framework) to compartmentalize it / prevent cross-tenant info leaks.

But let me rephrase what I was saying before:

> there's a way to take many transcripts of inference over a period, and convert/distil them together into an incremental-update training dataset (for memory, not for RLHF), that a model can be fine-tuned on as an offline batch process every day/week, such that a new version of the model can come out daily/weekly that hard-remembers everything you told it

As:

> for a given tenant user, there's a way to take all of their inference transcripts over a given period, and convert/distil them together into an incremental-update training dataset (for memory, not for RLHF), that a LoRA can be rebuilt (or itself fine-tuned) on. And that the work of all of these per-tenant LoRA rebuilds can occur asynchronously / "offline", on a batch-processing training cluster, gradually over the course of the day/week; such that at least once per day/week (presuming the tenant-user has any updated data to ingest), each tenant-user will get the effect of their own memory-LoRA being swapped out for a newer one.

---

Note how this is essentially what Apple claimed they would be doing with Apple Intelligence, re: "personal context."

The idea (that I don't think has ever come to fruition as stated—correct me if I'm wrong?) is that Apple would:

1. have your macOS and iOS devices spend some of their idle-on-charge CPU power to extract and normalize training fulltexts from whatever would be considered the user's "documents" — notes, emails, photos, maybe random text files on disk, etc.; and shove these fulltexts into some kind of iCloud-persisted database, where the fulltexts are PKI-encrypted such that only Apple's Private Compute Cloud (PCC) can decode them;

2. have the PCC produce a new/updated memory LoRA (or rather, six of them, because they need to separately imbue each of their domain-specific model "adapter" LoRAs with your personal-context memories);

3. and, once ready, have all your iCloud-account-synced devices to download the new versions of these memory-imbued adapter LoRAs.

---

And this is actually unnecessarily complex/circuitous for a cloud-hosted chat model. The ChatGPT/Claude/etc version of this architecture could be far simpler.

For a cloud-hosted chat model, you don't need a local agent to extract context from your devices; the context is just "past cloud-persisted chat transcripts." (But if you want "personal context" in the model, you could still get it, via an OpenClaw-style "personal agent"; such agents already essentially eat your files and spit them out external memories/RAGs/etc; the only change would be spitting them out into plain-old hidden-session chat transcripts instead, so as to influence the memories of the model they're running on.)

And you don't need a special securely-oblivious cluster to process that data, since unlike "Apple looking at the data on your computer" (which would upset literally everybody), nobody has any kind of expectation that e.g. OpenAI staff can't look at your ChatGPT conversation transcripts.

And cloud-hosted chat models don't really "do" domain-specific adapters (thus the whole "GPT" thing); so you only need to train one memory-LoRA per model. (Though I suppose that might still lead to training several LoRAs per user, if you're relying on smart routing to different models within a model family to save costs.)

And you don't need to distribute the memory-LoRAs back to client devices; as they can just live in an object store and get just-in-time loaded by the inference framework on a given node at the moment it begins an inference token-emission loop for a specific user. (Which might thus cause the inference cluster's routing to benefit from sticky sessions in a way it didn't before—but you don't need it; the LoRAs would likely be small enough to fetch and load within the ~second of delay it takes these cloud-hosted models to allocate you a node.)

Re: Learning from context is harder than we thought

#110

Earlier quoted context omitted.

That's not how training works - adjusting model weights to memorize a single data item is not going to fly. Model weights store abilities, not facts - generally. Unless the fact is very widely used and widely known, with a ton of context around it. The model can learn the day JFK died because there are millions of sparse examples of how that information exists in the world, but when you're working on a problem, you m…

I suspect we're going to need hypernetworks of some sort - dynamically generated weights, with the hypernet weights getting the dream-like reconsolidation and mapping into the model at large, and layers or entire experts generated from the hypernets on the fly, a degree removed from the direct-from-weights inference being done now. I've been following some of the token-free latent reasoning and other discussions arou…

Can I subscribe to your newsletter? You seem to be pretty plugged in to current research.
Post reply on HN