Is there some accessible explainer for what these numbers that keep going up actually mean? What happens at 100% accuracy or win rate?
Claude Sonnet 4.5
121–130 of 819 posts
Re: Claude Sonnet 4.5
#122Can't use Anthropic models in Cursor. Completely cost prohibitive compared to gpt-5 and grok models. Why is this? Does Anthropic have just higher infrastructure costs compared to OpenAI/xAI?
Re: Claude Sonnet 4.5
#123Is there some accessible explainer for what these numbers that keep going up actually mean? What happens at 100% accuracy or win rate?
edit: as far as what the numbers mean, they are arbitrary. They are only useful insofar as you can run two models (or two versions of the same model) on the same benchmark, and compare the numbers. But on an absolute scale the numbers don't mean anything.
Re: Claude Sonnet 4.5
#124Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
Is it number of lines? Tickets closed? PRs opened or merged? Number of happy customers?
Re: Claude Sonnet 4.5
#125Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
Re: Claude Sonnet 4.5
#126Earlier quoted context omitted.
More like churning benchmarks... Release new model at max power, get all the benchmark glory, silently reduce model capability in the following weeks, repeat by releasing newer, smarter model.
That (thankfully) can't compound, so would never be more than a one time offset. E.g. if you report a score of 60% SWE-bench verified for new model A, dumb A down to score 50%, and report a 20% improvement over A with new model B then it's pretty obvious when your last two model blogposts say 60%. The only way around this is to never report on the same benchmark versions twice, which they include too many to realisti…
Re: Claude Sonnet 4.5
#127Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
That's... not super surprising? SwiftUI changes pretty dang often, and the knowledge cutoff doesn't progress fast enough to cover every use-case.
I use Claude to write GTK interfaces, which is a UI library with a much slower update cadence. LLMs seem to have a pretty easy time working with bog-standard libraries that don't make giant idiomatic changes.
Re: Claude Sonnet 4.5
#128Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…
How do you measure 3x sustained output increase? Is it number of lines? Tickets closed? PRs opened or merged? Number of happy customers?
Have you heard of that study that shows AI actually makes developers less productive, but they think it makes them more productive??
EDIT: sorry all, I was being sarcastic in the above, which isn't ideal. Just annoyed because that "study" was catnip to people who already hated AI, and they (over-) cite it constantly as "evidence" supporting their preexisting bias against AI.
Re: Claude Sonnet 4.5
#129Earlier quoted context omitted.
I always wonder how absolute in performance a given model is. Sometimes i ask for Claude-Opus and the responses i get back are worse than the lowest end models of other assistants. Other times it surprises me and is clearly best in class. Sometimes in between this variability of performance it pops up a little survey. "How's Claude doing this session from 1-5? 5 being great." and i suspect i'm in some experiment of e…
> It would make sense they scale up and down depending on utilization right? It would, but > To state it plainly: We never reduce model quality due to demand, time of day, or server load. https://www.anthropic.com/engineering/a-postmortem-of-three-... If you believe them or not is another matter, but that's what they themselves say.
After all, using a different context window, subbing in a differently quantized model, throttling response length, rate limiting features aren’t technically “reducing model quality”.
Re: Claude Sonnet 4.5
#130Please y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT t…