Earlier quoted context omitted.
> My model doesn't need to know a single line of python. If I had to guess, the weights necessary to encode "how to program" are much larger than the final step of "output python."
ie what everyone asking for this fails to immediately realize.
Models Are Getting Dumber on Purpose
111–120 of 197 posts
Re: Models Are Getting Dumber on Purpose
#112Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
I think the vast majority of people do want general purpose models. They want to be able to ask it any question, or ask it to perform any task, and for it to do a decent job at it.
I agree that it's really hard (maybe even impossible) to build something that's everything for everyone. But your average (or even above-average) LLM user doesn't want to choose from a catalog to stitch together a model that does just what they need.
I do think for certain domains this is useful and will make sense: the model backing a coding harness doesn't need to know about the politics of 400BCE Rome. But I'm skeptical that many software developers will want to do what you propose, picking knowledge bases that are tailored to their current task or project. And at any rate, for web-based chat interfaces, most users just want to type a query and get an answer.
Re: Models Are Getting Dumber on Purpose
#113Earlier quoted context omitted.
This is a fundamental misunderstanding of how LLMs work. You can’t really specialize a model. You specialize the harness. A well-trained general purpose LLM doesn’t need examples in its training data, it can write good code in a new language you invented yesterday with just a spec definition. And it will perform better than a small model trained on lots of examples of your invented language. The reason is because of…
This is so right. We training Whisper Large model on 20,000 audio samples specific to a domain and it ended up reducing the ASR by 5% while improving WER of the finetuned domain by 0.5%. Instead we ended up with no finetuning. We give audio snippet to 2 AsR models, take 3 best transcriptions and ask the LLm to pick the best based on the context. That produced significantly higher accuracy in how an agent understands…
Re: Models Are Getting Dumber on Purpose
#114Earlier quoted context omitted.
Only if you fundamentally misunderstand what I am describing
You cannot have a single purpose LLM. Every single topic contributes to the performance of an LLM on any given area. You cannot have what you want. LLMs do not work like professions, college degrees or people. Using your examples, it is damaging to just know a single programming language, because there are patterns that are more common in say python than in swift, even though the only thing you want is swift, so LLM…
> so LLM performance in swift benefits from pythonic patterns
What I hear you saying is that the best way to make a swift-trained-only LLM smarter is to train it on some python too. And then with an infinite parameter budget, every other programming language or really any other data you train it on makes the model smarter - I accept that premise.
But in a fixed parameter budget, what is better? training on 50% Swift + 50% Python, or 50% Swift + 50% Rust. Because if I am doing Swift programming, I want whichever of the two is better for Swift. If I am doing Rust programming, maybe I want the model trained on 50% Python + 50% Rust. Sure, it would be smarter if you tossed in the swift code too - but we have a budget to stick to.
Now is it possible to make those pluggable? i.e. can you take a model trained on 50% python, and layer on 50% rust OR swift depending on what language you're using? Probably not right now, but maybe one day?
Re: Models Are Getting Dumber on Purpose
#115This makes hallucination detection more important. There's no reason that an LLM should have a vast number of obscure facts encoded. It can go out to a search engine for such facts. But the LLM has to be clear on what it doesn't know. (Google's pricing for search from programs starts at $2.50 per 1,000 queries. If an LLM reaches out to Google, it has to pay.)
The article really just glosses over this. Just because a model doesn’t kno a fact does not mean it will go and fetch it, the model just as well can invent the fact. For that to not happen the model needs to know that it’s missing the information and I don’t see the article making any explanation how this improves. I wonder what’s tre latest in this field? Did we get a grip on this problem?
That's the right question to ask. For a while, it seemed that hallucinations went down as models got bigger. That may only have been because, with a big enough model, the desired data might be in the model, somewhere, which would keep the model from making up something. That's the brute-force approach to the problem.
This new article indicates that trimming down the model by pulling out seldom used info makes the problem worse again.
If LLMs had reliable "I don't know", and access to search engines, much smaller models might work.
Re: Models Are Getting Dumber on Purpose
#116Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
> So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Tell me you don’t know how llm work without telling me you don’t know how llm work. That’s not how they work!
And then ideally, make it pluggable so I can pick what I want from off the shelf components, but if that's not possible - then just train up as many variants as you can so we can all pick the best variant for our current need.
Re: Models Are Getting Dumber on Purpose
#117Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
Re: Models Are Getting Dumber on Purpose
#118Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese p…
Re: Models Are Getting Dumber on Purpose
#119Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese p…
Yes, all three together would be even better. But it wouldn’t be if you had 100x more fanfiction, mostly synthetic, generated during RL to teach a model to be better at writing fan fiction. There are real limits to the amount of knowledge you can cram into fixed-size (downstream of hardware availability) weights. For a period scaling with data was basically “free” because we had the Internet and all the books/media that humans had already created; the data was accessible and limited (at least, the parts we think models should know about) enough and top-hardware big enough that we could basically compress the whole thing.
Post-training/RL are making this obsolete because they’re more about skill/capability acquisition rather than knowledge. They can generate much more data (most of it quotidian/useless, ie an agent made a typo in batch 382829) and clearly seem to cause a kind of mode collapse even in the most advanced frontier models.
We don’t need to make LLMs forget about SpongeBob SquarePants so they learn more about bash. But if I have a question about SpongeBob SquarePants, I don’t need to hear about load bearing seams prefaced with honest caveats after a model writes 400 lines of bash to look up SpongeBob’s family.
And there is probably a lot more SpongeBob knowledge we could put into models if we wanted to: interviews with the creative staff, a SpongeEnv/SpongeHarness modeling how the art/story team work together to create entertaining kids tv, a SpongeBench measuring entertainment value, etc. If a SpongeAgent spends 2000 years in Agent University learning how to Spongemaxx we probably don’t need or want to have it spend another 2000 years writing smoke tests
Re: Models Are Getting Dumber on Purpose
#120Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…
Sounds like unix philosophy. Or like Neo downloading Kung-Fu. Good either way.
https://linux.die.net/man/1/ls