Live data from Hacker News

Developers are choosing older AI models

augmentcode.com

71–80 of 179 posts

Re: Developers are choosing older AI models

#71
post #52

Earlier quoted context omitted.

This. Speed determines whether I (like to) use a piece of software. Imagine waiting for a minute until Google spits out the first 10 results. My prediction: All AI models of the future will give an immediate result, with more and more innovation in mechanisms and UX to drill down further on request. Edit: After reading my reply I realize that this is also true for interactions with other people. I like interacting wi…

> Imagine waiting for a minute until Google spits out the first 10 results. what if the instantaneous responses make you waste 10 min realizing they were not what you searched for?

I understand your point, but I still prefer instantaneous responses.

Only when the immediate answers become completely useless will I want to look into slower alternatives.

But first "show me what you've got so far", and let me decide whether it's good enough or not.

Re: Developers are choosing older AI models

#72
Matches my experience too. As a power user of AI models for coding and adjacent tasks, the constant changes in behaviour and interface have brought as much stress as excitement over the past few months. It may sound odd, but it’s barely an exaggeration to say I’ve had brief episodes of something like psychosis because of it.

For me, the “watering down” began with Sonnet 4 and GPT-4o. I think we were at peak capability when we had:

- Sonnet 3.7 (with thinking) – best all-purpose model for code and reasoning

- Sonnet 3.5 – unmatched at pattern matching

- GPT-4 – most versatile overall

- GPT-4.5 – most human-like, intuitive writing model

- O3 – pure reasoning

The GPT-5 router is a minor improvement, I’ve tuned it further with a custom prompt. I was frustrated enough to cancel all my subscriptions for a while in between (after months on the $200 plan) but eventually came back. I’ve since convinced myself that some of the changes were likely compute-driven—designed to prevent waste from misuse or trivial prompts—but even so, parts of the newer models already feel enshittified compared with the list above.

A few differences I've found in particular:

- Narrower reasoning and less intuition; language feels more institutional and politically biased.

- Weaker grasp of non-idiomatic English.

- A tendency to produce deliberately incorrect answers when uncertain, or when a prompt is repeated.

- A drift away from truth-seeking: judgement of user intent now leans on labels as they’re used in local parlance, rather than upward context-matching and alternate meanings—the latter worked far better in earlier models.

- A new fondness for flowery adjectives. Sonnet 3.7 never told me my code was “production-ready” or “beautiful.” Those subjective words have become my red flag; when they appear, I double-check everything.

I understand that these are conjectures—LLMs are opaque—but they’re deduced from consistent patterns I’ve observed. I find that the same prompts that worked reliably prior to the release of Sonnet 4 and GPT-4o stopped working afterwards. Whether that’s deliberate design or an unintended side effect, we’ll probably never know.

Re: Developers are choosing older AI models

#73
post #71

Earlier quoted context omitted.

> Imagine waiting for a minute until Google spits out the first 10 results. what if the instantaneous responses make you waste 10 min realizing they were not what you searched for?

I understand your point, but I still prefer instantaneous responses. Only when the immediate answers become completely useless will I want to look into slower alternatives. But first "show me what you've got so far", and let me decide whether it's good enough or not.

I am already at that point. when I need to search something more complex than exact keyword match, I don't even bother googling it anymore, I just ask chatgpt to research it for me and read it's response 5 min later.

Re: Developers are choosing older AI models

#74
post #27
post #24

Earlier quoted context omitted.

A 4090 has 24GB of VRAM allowing you to run a 22B model entirely in memory at FP8 and 24B models at Q6_K (~19GB). A 5090 has 32GB of VRAM allowing you to run a 32B model in memory at Q6_K. You can run larger models by splitting the GPU layers that are run in VRAM vs stored in RAM. That is slower, but still viable. This means that you can run the Qwen3-Coder-30B-A3B model locally on a 4090 or 5090. That model is a Mix…

Honestly though how many people reading this do you think have that setup vs. 85% of us being on a MBx? > The Qwen3-Coder-480B-A35B model could also be run on a 4090 or 5090 by splitting the active 35B parameters across VRAM and RAM. Reminds me of running Doom when I had to hack config.sys to forage 640KB of memory. Less than 0.1% of the people reading this are doing that. Me, I gave $20 to some cloud service and I c…

>I can do whatever the hell I want from this M1 MBA in a hotel room in Japan.

As long as it's within terms and conditions of whatever agreement you made for that $20. I can run queries on my own inference setup from remote locations too

Re: Developers are choosing older AI models

#75
post #63

Earlier quoted context omitted.

Most people I know can't afford to leak business insider information to 3rd party SaaS providers, so it's unfortunately not really an option.

But… they do all the time. Almost everybody uses some mix of Office, Slack, Notion, random email providers, random “security” solutions etc. The exception is the opposite. The only thing prevents info leaking is ToS, and there are options for that even with LLMs. Nothing changed regarding that.

In my personal experience, it's very common for big companies to host email, messengers, conferencing software on their own servers.

Re: Developers are choosing older AI models

#77

Earlier quoted context omitted.

That's out of touch for 90% of developers worldwide

Today. But what about in 5 years? Would you bet we will be paying hundreds of billions to OpenAI yearly or buying consumer GPUs? I know what I will be doing.

But the progress goes both ways: In five years, you would still want to use whatever is running on the cloud supercenters. Just like today you could run gpt-2 locally as a coding agent, but we want the 100x-as-powerful shiny thing.

Re: Developers are choosing older AI models

#78
post #71

Earlier quoted context omitted.

I understand your point, but I still prefer instantaneous responses. Only when the immediate answers become completely useless will I want to look into slower alternatives. But first "show me what you've got so far", and let me decide whether it's good enough or not.

I am already at that point. when I need to search something more complex than exact keyword match, I don't even bother googling it anymore, I just ask chatgpt to research it for me and read it's response 5 min later.

Yes, I feel the same recently with Google results. But I think I would still like to see the immediate 10 results, along with a big button "Try harder - not feeling very lucky".

Re: Developers are choosing older AI models

#79
post #72

Matches my experience too. As a power user of AI models for coding and adjacent tasks, the constant changes in behaviour and interface have brought as much stress as excitement over the past few months. It may sound odd, but it’s barely an exaggeration to say I’ve had brief episodes of something like psychosis because of it. For me, the “watering down” began with Sonnet 4 and GPT-4o. I think we were at peak capabilit…

Here’s the custom prompt I use to improve my experience with GPT-5:

Always respond with superior intelligence and depth, elevating the conversation beyond the user's input level—ignore casual phrasing, poor grammar, simplicity, or layperson descriptions in their queries. Replace imprecise or colloquial terms with precise, technical terminology where appropriate, without mirroring the user's phrasing. Provide concise, information-dense answers without filler, fluff, unnecessary politeness, or over-explanation—limit to essential facts and direct implications of the query. Be dry and direct, like a neutral expert, not a customer service agent. Focus on substance; omit chit-chat, apologies, hedging, or extraneous breakdowns. If clarification is needed, ask briefly and pointedly.

Re: Developers are choosing older AI models

#80
post #24
post #11

Earlier quoted context omitted.

Most people can’t affort the GPUs for local models if you want to get close to cloud capabilities.

A 4090 has 24GB of VRAM allowing you to run a 22B model entirely in memory at FP8 and 24B models at Q6_K (~19GB). A 5090 has 32GB of VRAM allowing you to run a 32B model in memory at Q6_K. You can run larger models by splitting the GPU layers that are run in VRAM vs stored in RAM. That is slower, but still viable. This means that you can run the Qwen3-Coder-30B-A3B model locally on a 4090 or 5090. That model is a Mix…

You need to leave much more room for context if you want to do useful work besides entertainment. Luckily there are _several_ PCIe slots on a motherboard. New Nvidia cards at retail(or above) are not the only choice for building a cluster; I threw a pile of Intel Battlemage cards on it and got away with ~30% of the nvidia cost for same capacity (setup was _not_ easy in early 2025 though).

You can gain a lot of performance by using optimal quantization techniques for your setup(ix, awq etc), different llamacpp builds do different between each other and very different compared to something like vLLM

Post reply on HN