Earlier quoted context omitted.
I'm saying that the kind of changes you propose aren't made by anyone, and might generally not be worth making. Because "better RLVR" is an easier and better pathway to actual cross-domain performance gains. If you could stabilize the kind of mess you want to make, you could put that effort into better RL objectives and get more return.
The mainstream LLM crowd aren't making these sorts of major changes yet, although some like DeepMind (the OG pushers of RL for AGI!) do acknowledge that a few more "transformer level" breakthoughs are necessary to reach what THEY are calling AGI, and others like LeCun are calling for more animal-like architectures. Anyways, regardless of who is currently trying to move beyond LLMs or not, it should be pretty obvious…
"Fundamental limitations" aren't actually fundamental. If you want more learning than what "in-context" gives you? Teach the usual "CLI agent" LLM to make its own LoRAs and there goes that. So far, this isn't a bottleneck so pressing you'd want to resolve it by force.
LeCun is laughing stock nowadays, he didn't get kicked out of Meta for no reason.