Live data from Hacker News

Developers are choosing older AI models

augmentcode.com

171–179 of 179 posts

Re: Developers are choosing older AI models

#171

Earlier quoted context omitted.

> There may be healthy margins for them but I would bet it’s always going to be relatively cheaper for me to pay for the compute rather than manage it myself. Depends almost completely on usage. No one is renting out hardware 24x7 and making a loss on it. If you only have sporadic use then renting is better. If you're running it almost all the time of purchasing it outright is better.

Sure but we were talking about gaming rigs to run models locally. You are describing some extreme edge folks that are keeping 24/7 work on gaming rigs in your home.

> Sure but we were talking about gaming rigs to run models locally. You are describing some extreme edge folks that are keeping 24/7 work on gaming rigs in your home.

In that scenario the case is even weaker for the rented-hardware model - if you're going to have a gaming rig, you're only paying a little bit more on top for a GPU with more RAM, not the full cost of the rig.

The comparison then is the extra cost of using a 24GB GPU over a standard gaming rig GPU (12GB? 8GB?) versus the cost of renting the GPU whenever you need it.

Re: Developers are choosing older AI models

#172

Earlier quoted context omitted.

Sure but we were talking about gaming rigs to run models locally. You are describing some extreme edge folks that are keeping 24/7 work on gaming rigs in your home.

> Sure but we were talking about gaming rigs to run models locally. You are describing some extreme edge folks that are keeping 24/7 work on gaming rigs in your home. In that scenario the case is even weaker for the rented-hardware model - if you're going to have a gaming rig, you're only paying a little bit more on top for a GPU with more RAM, not the full cost of the rig. The comparison then is the extra cost of us…

Honestly not sure what you are talking about.

I could either spend $20 a month for my cursor license.

Or

Spend $2k+ upfront to build a machine to run models locally. Pay for the electricity cost and time to set both the machine and software up.

Re: Developers are choosing older AI models

#173

Earlier quoted context omitted.

> Sure but we were talking about gaming rigs to run models locally. You are describing some extreme edge folks that are keeping 24/7 work on gaming rigs in your home. In that scenario the case is even weaker for the rented-hardware model - if you're going to have a gaming rig, you're only paying a little bit more on top for a GPU with more RAM, not the full cost of the rig. The comparison then is the extra cost of us…

Honestly not sure what you are talking about. I could either spend $20 a month for my cursor license. Or Spend $2k+ upfront to build a machine to run models locally. Pay for the electricity cost and time to set both the machine and software up.

> Spend $2k+ upfront to build a machine to run models locally.

You said this was in the context of a gaming rig. You're not spending an extra $2k on your gaming rig to run models locally.

If you're building a dedicated LLM machine OR you're using less compute than you are paying the provider for, then, yup - $20/m is cheaper.

When you start using the model more, or if you're already building a gaming rig, then it's going to be cheaper to self-host.

Re: Developers are choosing older AI models

#174

Earlier quoted context omitted.

Yeah, I'm just going through the Cerebras migration at the moment. It's a shame Cerebras completely dropped Qwen3 Coder's fast tool calling, short and instant responses, and better speed overall for GLM 4.6 thinking. Qwen3 is really good at hitting the tools first, then coming up with a well-grounded answer based on reality. Sometimes it's good when a model is Socratic: just knows it knows nothing. GLM 4.6 on the oth…

> Qwen3 is really good at hitting the tools first, then coming up with a well-grounded answer based on reality. I don't know, I had a lot of issues with Qwen models when it comes to RooCode/Cline - failed edits (albeit with a requirement for 100% precision, since I don't want the wrong lines to be replaced) or calling tools without parameters (e.g. list_files without path) and also stuff like using wrong path separat…

I've used it with CC and the match was great, not a lot of issues, I believe Qwen had a clear focus on distilling Anthropic models. GLM 4.6 is slightly better maybe, but the speed dropped to half on Cerebras so that's the price for maybe ~15% improvement in model overall quality. This quality does not necessarily means the end product (the code) is 15% better, just that now I take 12 turns with GLM instead of 15 turns with Qwen to get something done, but turn speed has been reduced to half in Cerebras, so my TTC (time-to-completion) has actually gone from 15min to 24min!

Re: Developers are choosing older AI models

#175
post #28

Earlier quoted context omitted.

I have canceled my Claude Max subscription because Sonnet 4.5 is just too unreliable. For the rest of the month I'm using Opus 4.1 which is much better but seems to have much lower usage limits than before Sonnet 4.5 was released. When I hit 4.1 Opus limits I'm using Codex. I will probably go through with the Codex pro subscription.

> [...] I'm using Opus 4.1 which is much better but seems to have much lower usage limits than before Sonnet 4.5 was released [...] Yes, it's down from 40h/week to 3-5h/week on Max plan, effectively. A real bummer. See my comment here [1] regarding [2]. [1] https://news.ycombinator.com/item?id=45604301 [2] https://github.com/anthropics/claude-code/issues/8449

Thanks, didn't know that but aligns with my experience

Re: Developers are choosing older AI models

#176

Earlier quoted context omitted.

Honestly not sure what you are talking about. I could either spend $20 a month for my cursor license. Or Spend $2k+ upfront to build a machine to run models locally. Pay for the electricity cost and time to set both the machine and software up.

> Spend $2k+ upfront to build a machine to run models locally. You said this was in the context of a gaming rig. You're not spending an extra $2k on your gaming rig to run models locally. If you're building a dedicated LLM machine OR you're using less compute than you are paying the provider for, then, yup - $20/m is cheaper. When you start using the model more, or if you're already building a gaming rig, then it's g…

I think the plot flew way over your head. We were comparing costs and for my reading, saying gaming rig is more about consumer grade hardware and not so much an assumption that you already have one. After all I assume we would be buying a 5090 for the vram and at current market price that’s $3k alone. You would probably end up spending at least $20 in electricity and cooling every month of you are running it near 24/7.

So again, the economics don’t really make sense except in specific edge cases or for folks that don’t want to pay vendors. Also please don’t use italics, I don’t know why but every time you see them used it’s always a silly comment.

Re: Developers are choosing older AI models

#177
post #69
post #22

To the authors of the site, please know that your current "Cookiebot by Usercentrics" is old and pretty much illegal. You shouldn't need to click 5 times to "Reject all" if accepting all is one click. Newer versions have a "Deny" button.

Just set up your browser to never even load that BS.

I cannot audit and report GDPR violations if I do that.

Re: Developers are choosing older AI models

#178

I wish we could pin down not only the model but also the way the UI works as well. Last week Claude seemed to have a shift in the way it works. The way it summarises and outputs its results is different. For me it's gotten worse. Slower, worse results, more confusing narrowing down what actually changed etc etc. Long story short, I wish I was able to checkpoint the entire system and just revert to how it was previous…

I agree, much slower and worse output. It is substantially worse now than it was weeks ago.

It spends a lot of time coming up with “UI options” (Select 1, 2 or 3 with a TUI interface) for me to consider when it could just ask me what I want, not come up with a 5 layer flow chart of possibilities.

Overall I think it is just Anthropic tweaking things to reduce costs.

I am paying for a Max subscription but I am going to reevaluate other options.

Re: Developers are choosing older AI models

#179
post #81

I've found that the VSCode GitHub Copilot extension defaults to Claude Sonnet 4.0 (in agent mode) in all new workspaces. It's the first thing I check now, but I imagine a lot of people just roll with it, especially if they use inline completions where it might not be obvious what model is being used.

I've seen similar behavior, even after having selected 4.5
Post reply on HN