Live data from Hacker News

Models Are Getting Dumber on Purpose

w4g1.dev

121–130 of 197 posts

Re: Models Are Getting Dumber on Purpose

#121
post #117

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

Sounds a bit like 'I want to make horses faster, surely I won't need mechanical engineering knowledge'. We don't know everything that we don't know, so it's hard to say what we don't need to know.

But I am not asking for a faster horse. I am asking for a draft horse instead of a race horse. I know the tasks I have on hand, I know the VRAM and compute budget I have to run them. I am not asking for AGI, I just want something to edit my little text files.

Re: Models Are Getting Dumber on Purpose

#122

Earlier quoted context omitted.

ie what everyone asking for this fails to immediately realize.

You're a bingo. It's obvious that a model could know no Python, since Python could not exist in the world in the first place.

the best llm for coding is the one that knows every programming language imagined by a caffeine-fuelled comp-sci student at 2am but never built.

Re: Models Are Getting Dumber on Purpose

#123

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese p…

People keep forgetting that programming is not just about knowing the target programming language, but also an enormous volume of tacit knowledge:

    - Understanding of protocols like HTTP.
    - HTML, JS, CSS, SVG, and everything "web".
    - Understanding of databases, SQL, etc.
    - Abstract code architecture patterns.
    - Understanding the users' requests in English.
    - Responding in English.
    - Command line tool usage (agents/harnesses)
    - Industry-specific knowledge that can be applied.
    - Frameworks, SDKs, applicable libraries.
    - Relevant legal requirements.
    - Etc...
I.e.: If I tell a frontier AI that this project is for a "local council in XYZ location" it can immediately figure out that a scalable, globally distributed architecture is not required. It can also figure out that using local time instead of UTC is not only "fine", but even desired. Or that globalization/localization is not required... or.... required if the council is in some place like Belgium or Canada where multiple languages are officially recognised and supported by the government.

Re: Models Are Getting Dumber on Purpose

#124
post #112

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

> I don't want general purpose models. They try to be everything to everyone. I think the vast majority of people do want general purpose models. They want to be able to ask it any question, or ask it to perform any task, and for it to do a decent job at it. I agree that it's really hard (maybe even impossible) to build something that's everything for everyone. But your average (or even above-average) LLM user doesn'…

I want to pick it myself because I want to run this stuff locally on hardware I can afford today. But most people seem happy enough using the cloud, where this kind of architecture could be seamless if it existed. I.e. the query to `/chat/completions` contains an extra parameter `domains: ["c++", "swift", "navigation" "gis"]` and then those are the modules that get allocated to this query. When you start a new conversation no domains are set and it hits a generalist model, but the generalist model includes domains so the next query doesn't have to hit the generalist model. The model could even have a tool it could call to rope in new domains if the scope expanded to include other modules.

Re: Models Are Getting Dumber on Purpose

#125

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

This just means you have to describe what you want in Swift or whatever. If it doesn’t have the language then it doesn’t have the capability to transform intent into code.

Re: Models Are Getting Dumber on Purpose

#126

Earlier quoted context omitted.

You're a bingo. It's obvious that a model could know no Python, since Python could not exist in the world in the first place.

the best llm for coding is the one that knows every programming language imagined by a caffeine-fuelled comp-sci student at 2am but never built.

if neuralink ever becomes a thing, thoughts about programming might be stolen for LLM training data lol

Re: Models Are Getting Dumber on Purpose

#127
post #117

Earlier quoted context omitted.

Sounds a bit like 'I want to make horses faster, surely I won't need mechanical engineering knowledge'. We don't know everything that we don't know, so it's hard to say what we don't need to know.

But I am not asking for a faster horse. I am asking for a draft horse instead of a race horse. I know the tasks I have on hand, I know the VRAM and compute budget I have to run them. I am not asking for AGI, I just want something to edit my little text files.

At a more technical level, what do you suggest? Training a small LLM on Python code exclusively? And then one on general CS/algorithms, which you'll also need? I don't think the current transformer architectures would compose as you suggest.

Re: Models Are Getting Dumber on Purpose

#128

Earlier quoted context omitted.

An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese p…

People keep forgetting that programming is not just about knowing the target programming language, but also an enormous volume of tacit knowledge : - Understanding of protocols like HTTP. - HTML, JS, CSS, SVG, and everything "web". - Understanding of databases, SQL, etc. - Abstract code architecture patterns. - Understanding the users' requests in English. - Responding in English. - Command line tool usage (agents/ha…

Those assumptions are just that - assumptions. "Local council in XYZ location" implies a bunch of things, and each one might be wrong for my specific circumstances. What better way to guide expectations than importing specific knowledge? I.e. if I import the english and catalan modules, then I probably want to localize my site in english and catalan.

It would be trivial to have a pre-flight convo with an llm to guide the user thru module choices. "Build a site" -> "ok, describe the purpose" -> "local council in XYZ location" -> "that implies you won't need localization since XYZ has a monolingual government" -> "english and catalan localization please".

Right now, you prompt and it builds using assumptions, and we prompt to adjust. I think it would be great to be able to pre-load a set of assumptions.

Re: Models Are Getting Dumber on Purpose

#129
post #127

Earlier quoted context omitted.

But I am not asking for a faster horse. I am asking for a draft horse instead of a race horse. I know the tasks I have on hand, I know the VRAM and compute budget I have to run them. I am not asking for AGI, I just want something to edit my little text files.

At a more technical level, what do you suggest? Training a small LLM on Python code exclusively? And then one on general CS/algorithms, which you'll also need? I don't think the current transformer architectures would compose as you suggest.

With what I know about how LLMs work now, I guess I am suggesting more specific variants. Qwen3.8 has a 2.4T version and a 27B version. I understand that to mean that they are the same architecture, just one version has a massive training set and the other has a very small subset. So, it seems very possible that variants of 27B could be generated that tune it for specific things by selecting different training data from the large corpus. One model for Python, another for Swift, another for research, another for creative writing, etc.

I think you're right that current architectures don't compose like that - but I feel like that's a result of the focus on MOAR DATA, and a "race for AGI" - if we set those ideas aside, a more composable architecture seems very possible.

Re: Models Are Getting Dumber on Purpose

#130

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

This is a fundamental misunderstanding of how LLMs work. You can’t really specialize a model. You specialize the harness. A well-trained general purpose LLM doesn’t need examples in its training data, it can write good code in a new language you invented yesterday with just a spec definition. And it will perform better than a small model trained on lots of examples of your invented language. The reason is because of…

Unless you are building a factory assembly line where a model is literally doing the same thing over and over, you almost always want a general purpose model over a specialized one.

Turns out the world is made of simple, specialist processes, not generalists trying to achieve them. Adaptability may be of great benefit in evolutionary terms or for a walking anthropoid, but the majority of biology, chemistry, and mathematics rely upon specialist process for good reason. See also the old trope about robotics: that's what you call it before it works, otherwise it'd be a dishwasher.

The upshot is: use a generalist to create a simple solution once, and scale that. Don't deploy the generalist at scale, that's a waste of resources and an inefficient solution.

Post reply on HN