I feel that the recent iterations of LLM haven't provided an intuitive qualitative leap. Have they entered a bottleneck period so quickly?
Azure recently discontinued the gpt-4.1 model. I had to move off of this model, and moving to any gpt-5* model was worse (higher failures & less accuracy), and more expensive. I had to rewrite the entire system from high school level prompts to lower elementary school level prompts using non-gpt models. I would say models entered a bottleneck a long time ago. My personal opinion is now they are overfitting newer mode…
GPT-5.5 Price Increase: What It Costs
61–70 of 80 posts
Re: GPT-5.5 Price Increase: What It Costs
#62Re: GPT-5.5 Price Increase: What It Costs
#63Earlier quoted context omitted.
Yes, the signal we are measuring is quite different from most evals. We are measuring sth much closer to: when multiple agents compete on the same spec, which one produces the patch that holds up best in code review? Most evals are static / synthetic, and for code, generally stop at tests. Test evals are weak proxies for quality since it's difficult to encode qualities like scope creep/churn, codebase fit, maintainab…
Ok, but my point is that the claims you make about more reasoning performing worse seems kinda suspicious and I haven't seen any analysis exploring why that would happen.
Re: GPT-5.5 Price Increase: What It Costs
#64Earlier quoted context omitted.
Ok, but my point is that the claims you make about more reasoning performing worse seems kinda suspicious and I haven't seen any analysis exploring why that would happen.
My point is more reasoning often leads to worse "scope creep/churn, codebase fit, maintainability".
Re: GPT-5.5 Price Increase: What It Costs
#65Earlier quoted context omitted.
My point is more reasoning often leads to worse "scope creep/churn, codebase fit, maintainability".
I get it, but that is a significant claim. And the claim could be right, but it could also be wrong, and I see no analysis, not even a blog post on your website saying "wow, look at this weird thing we found". To me that makes the claim suspicious because it signals that nobody thought to investigate what's going on. Investigating weird results is how we demonstrate that what we're doing is right.
We are not the only ones to see the reasoning inversion.: https://arxiv.org/abs/2510.11977, https://arxiv.org/abs/2502.08235, https://arxiv.org/abs/2507.14417
Re: GPT-5.5 Price Increase: What It Costs
#66Earlier quoted context omitted.
They only did that after they "found" ~300k H100 equivalent compute. Before signing that deal they were severely compute constrained. Especially visible when EU tz was still active and US east would wake up.
I don't disagree. It's just a weird way to describe them currently when they just announced massively increasing limits.
Re: GPT-5.5 Price Increase: What It Costs
#67I feel that the recent iterations of LLM haven't provided an intuitive qualitative leap. Have they entered a bottleneck period so quickly?
Azure recently discontinued the gpt-4.1 model. I had to move off of this model, and moving to any gpt-5* model was worse (higher failures & less accuracy), and more expensive. I had to rewrite the entire system from high school level prompts to lower elementary school level prompts using non-gpt models. I would say models entered a bottleneck a long time ago. My personal opinion is now they are overfitting newer mode…
Re: GPT-5.5 Price Increase: What It Costs
#68Earlier quoted context omitted.
Azure recently discontinued the gpt-4.1 model. I had to move off of this model, and moving to any gpt-5* model was worse (higher failures & less accuracy), and more expensive. I had to rewrite the entire system from high school level prompts to lower elementary school level prompts using non-gpt models. I would say models entered a bottleneck a long time ago. My personal opinion is now they are overfitting newer mode…
I am wondering if everyone is moving to an IPO and striking these bizarre circular deals because they’ve hit the ceiling on what can be done with more compute until a major architectural innovation happens. Still amazing, but 5.5 does feel like incremental progress with a massive up charge.
The reality is both Anthropic and OAI have converged on LLMs as being a thing for software production - that's where the majority of their revenue is coming from.
Re: GPT-5.5 Price Increase: What It Costs
#69New model releases are now like new iPhones--mostly imperceivable improvements with a higher price tag. That's one of the major benefits to open source: you can "freeze" what model you're using. Often it's the model that you know that wins over the one that is different enough that you have to start from scratch with every major update. Most businesses require cost control and predictability over a cutting edge with…
If I skip 2 models of iPhone upgrades, there is definitely a difference between how the thing feels - and it feels its worth the money.
If I skip 2 models of upgrades of the frontier models now, I highly doubt I can discern what the difference is and what exactly I'd be paying more for.
Re: GPT-5.5 Price Increase: What It Costs
#70For personal use I switched to coding plans containing GLM 5.1, Kimi K2.6 and Xiaomi MiMo V2.5 Pro and I never been happier. I said goodbye to both Claude Max and Cursor.