I'm confused by how Opus is presented to be superior in nearly every way for coding purposes yet the general consensus and my own experience seem to be that Sonnet is much much better. Has anyone switched to entirely using Opus from Sonnet? Or maybe switching to Opus for certain things while using Sonnet for others?
Im on the Max plan and generally Opus seems to do better work than Sonnet. However, that’s only when they allow me to use Opus. The usage limits, even on the max plan, are a joke. Yesterday I hit the limits within MINUTES of starting my work day.
Claude Opus 4.1
41–50 of 344 posts
Re: Claude Opus 4.1
#42I'm confused by how Opus is presented to be superior in nearly every way for coding purposes yet the general consensus and my own experience seem to be that Sonnet is much much better. Has anyone switched to entirely using Opus from Sonnet? Or maybe switching to Opus for certain things while using Sonnet for others?
Given that there’s nothing close to scientific analysis going on, I find it hard to tell how big the “Sonnet is overall better, not just sometimes” crowd is. I think part of the problem is that “The bigger model is better” feels obvious to say, so why say it? Whereas “the smaller model is better actually” feels both like unobvious advice and also the kind of thing that feels smart to say, both of which would lead to more people who believe it saying it, possibly creating the illusion of consensus.
I was trying to dig into this yesterday, but every time I come across a new thread the things people are saying and the proportions saying what are different.
I suppose one useful takeaway is this: If you’re using Claude Max and get downgraded from Opus to Sonnet for a few hours, you don’t have to worry too much about it being a harsh downgrade in quality.
Re: Claude Opus 4.1
#43it is barely an improvement according to their own benchmarks. not saying thats a bad thing, but not enough for anybody to notice any difference
I don't think this could even be called an improvement? It's small enough that it could just be random chance
Instead, ideally they’d run the benchmark tests many times, and share all of the results so we could make statistical determinations.
Re: Claude Opus 4.1
#44Re: Claude Opus 4.1
#45Re: Claude Opus 4.1
#46Has anyone tested it yet? How's it acting?
Re: Claude Opus 4.1
#47it is barely an improvement according to their own benchmarks. not saying thats a bad thing, but not enough for anybody to notice any difference
That's why they named it 4.1 and not 4.5
Re: Claude Opus 4.1
#48At least Sonnet 4 is still usable, but I'll be honest, it's been producing worse and worse slob all day.
I've basically wasted the morning on Claude Code when I should've just been doing it all myself.
Re: Claude Opus 4.1
#49just ran the LLM to SQL benchmark over opus-4.1 and it didn't top previous version :thinking: => https://llm-benchmark.tinybird.live/
LLMs are non-deterministic, I think benchmarks should be more about averages of N runs, rather than single shot experiments.
Re: Claude Opus 4.1
#50I'm confused by how Opus is presented to be superior in nearly every way for coding purposes yet the general consensus and my own experience seem to be that Sonnet is much much better. Has anyone switched to entirely using Opus from Sonnet? Or maybe switching to Opus for certain things while using Sonnet for others?