Live data from Hacker News

Models Are Getting Dumber on Purpose

w4g1.dev

151–160 of 197 posts

Re: Models Are Getting Dumber on Purpose

#151

Earlier quoted context omitted.

Those assumptions are just that - assumptions. "Local council in XYZ location" implies a bunch of things, and each one might be wrong for my specific circumstances. What better way to guide expectations than importing specific knowledge? I.e. if I import the english and catalan modules, then I probably want to localize my site in english and catalan. It would be trivial to have a pre-flight convo with an llm to guide…

See: https://en.wikipedia.org/wiki/Bitter_lesson Everyone assumes that carefully crafting a specific AI architecture with bits and pieces bolted together based on their human intuition is necessarily superior to simply using a bigger monolithic AI model. It turns out that the opposite is true, and has been demonstrated over and over again. The bitter lesson is this: You can simply ask a frontier model to do the thing…

Wouldn't this imply that in terms of AI usage you should take a "Wait and see" approach? I.e. just wait until the models can easily do whatever it is you want?

Re: Models Are Getting Dumber on Purpose

#152

Earlier quoted context omitted.

could you share the 3blue1brown videos you're referring to?

https://youtu.be/l6DKRf-fAAM maybe? Title is “Reinventing Entropy” And of course the neural network series.

Yes, that's the one

Re: Models Are Getting Dumber on Purpose

#153
post #99

Earlier quoted context omitted.

I agree with everything you say except this: > You can make any modern LLM explain its reasoning You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".

Tangent: This is often true of humans as well. We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.

And we all know some people that rationalise poor choices and misbehavior, hide their mistakes, etc, to an unacceptable degree. Sometimes the individual knows they are rationalising but continues anyway, other times they seem incapable of seeing that.

When you ask people who are rationalising poor behaviour about the scenario, but it is someone else doing it, they may arrive at a better answer. Can we use multiple LLMs to achieve self criticism and critical thinking?

Re: Models Are Getting Dumber on Purpose

#154
post #39

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

I strongly suspect this will be the future

The future will be "AI operator license #123..., subclass 'Primitive Coding'", with your TouchID / FaceID enabled only. And I am not being sarcstic.

Re: Models Are Getting Dumber on Purpose

#155
post #144

This AI generated post (100% on Pangram) is pretty out of date. >On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questions. SimpleQA hasn't been updated in a long time. Gemini 2.5 Pro is a sixteen-month-old model, not "the best recall money can buy". >The part I find most promising is what this does t…

Yea the "When the fact lives outside the model, a wrong answer has an address" sentence seems aggressively AI written. Saw that and my senses went off.

Senses of what? LOL. The whole Internet is AI generated by now and we all contribute to that on daily basis. get used to it or dull your senses ...

Re: Models Are Getting Dumber on Purpose

#156

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

I shudder to imagine the monetization schemes of such an architecture though.

Re: Models Are Getting Dumber on Purpose

#157

Earlier quoted context omitted.

This is a fundamental misunderstanding of how LLMs work. You can’t really specialize a model. You specialize the harness. A well-trained general purpose LLM doesn’t need examples in its training data, it can write good code in a new language you invented yesterday with just a spec definition. And it will perform better than a small model trained on lots of examples of your invented language. The reason is because of…

This is so right. We training Whisper Large model on 20,000 audio samples specific to a domain and it ended up reducing the ASR by 5% while improving WER of the finetuned domain by 0.5%. Instead we ended up with no finetuning. We give audio snippet to 2 AsR models, take 3 best transcriptions and ask the LLm to pick the best based on the context. That produced significantly higher accuracy in how an agent understands…

Deep Fusion is best, when words and phrase patterns in the domain are known. Deep Fusion means to hint the Whisper decoder about the next possible words using LLM-in-the-loop.

Re: Models Are Getting Dumber on Purpose

#158

Ideally what I'd like to see is pluggable knowledge bases. So if I'm e.g. coding a SwiftUI app for navigation, I'd take 9B of basic coding and reasoning, add 10B of swift/swiftUI, add 5B of GIS/geography knowledge and another 5B of frontend app design knowledge. My model doesn't need to know a single line of python. Then when I want to research electronics components, I grab a 15B model of agentic research techniques…

An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant. The model finds a structure to the problem you give it in a largely language-agnostic way that benefits from training on every language (these things are direct descendants of Google Translate), and even non-programming knowledge - the structure of your task might resemble an ancient Chinese p…

An LLM works better the more disparate world knowledge it has, even if it's not immediately obvious why it would be relevant.

You just defined a liberal arts education.

Re: Models Are Getting Dumber on Purpose

#159
post #99

Earlier quoted context omitted.

I agree with everything you say except this: > You can make any modern LLM explain its reasoning You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".

Tangent: This is often true of humans as well. We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.

Tangent on the tangent: I think that's true in a minority of cases and in a majority of AI cases. Though in principle I think it should be possible for an LLM to have access to and faithfully represent its own reasoning.

Re: Models Are Getting Dumber on Purpose

#160
post #137

Earlier quoted context omitted.

See: https://en.wikipedia.org/wiki/Bitter_lesson Everyone assumes that carefully crafting a specific AI architecture with bits and pieces bolted together based on their human intuition is necessarily superior to simply using a bigger monolithic AI model. It turns out that the opposite is true, and has been demonstrated over and over again. The bitter lesson is this: You can simply ask a frontier model to do the thing…

>I get it. You don't feel ownership over someone else's AI. You don't feel involved, you don't feel like you have agency. You don't _have_ ownership of someone else's ai, and that comes with real risks. Security risks, privacy risks, business risk. They might rug pull you, they might charge you more, or like atrophic, silently corrupt the answers, or code... The labs are happy to jump on any emergent capability the s…

It doesn't have to be "externally hosted, proprietary AI model"! The argument is against "self-assembled small AI pieces" versus frontier monolithic models.

A) You can always self-host something like Kimi, DeepSeek, or GLM.

B) Just because you use a specific proprietary AI for programming doesn't actually bind you to that provider in any meaningful way. The authored code remains even if you stop paying them!

Of course, if you use AI as an active component in some sort of service, then the EULA, rug-pulls, etc... suddenly start to matter. That's a different story.

Post reply on HN