Live data from Hacker News

Claude Opus 5

anthropic.com

301–310 of 1001 posts

Re: Claude Opus 5

#301

Looking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

I don't think can use the AA index to say something is 10% smarter

I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1

Re: Claude Opus 5

#302

Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.

My understanding is that Opus should be used for planning, macro-level conversations and Sonnet for execution. So, for coding, for example: Opus for solution design and architectural blueprint and then Sonnet for actual implementation. Works out cheaper with minimal loss of quality. At least that's my personal understanding and anecdotal experience.

It depends on your quality bar. At a fixed level of quality, given a high reasoning sonnet vs a low reasoning opus, the low reasoning opus tends to be pareto optimal.

It's only when you need even lower levels of cost than opus at zero to low reasoning when sonnet starts to make sense at all.

Re: Claude Opus 5

#303
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

> coding is largely solved - Boris

The code it outputs, yes! It's fantastic. It's just so frustrating that the product and UX before the code output is so bad. Greatness is so close within their reach, if only they invested in product and QA people.

Re: Claude Opus 5

#304
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination

Re: Claude Opus 5

#305
post #90
post #40

I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.

I guess character? Fable is more friendly and curious while opus is a bit more deliberate and conservative.

Surely that's tunable? OpenAI lets you tune response characteristics.

Re: Claude Opus 5

#306

Earlier quoted context omitted.

Openrouter should ideally kill in this space and make their model agnostic infra like memory, harnesses, chat applications.

OpenRouter is in acquisition talks with Stripe, fyi I would expect routers to commodify like tokens.

OpenRouter sprang up overnight. I might need to replace some urls and access tokens should they decide to try and screw me.

Re: Claude Opus 5

#307

But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed. I've never trusted on model cards though. I'm sorry.

Another benefit is that fast mode can be used on subscription, but Anthropic's won't

Re: Claude Opus 5

#308
post #204

Earlier quoted context omitted.

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.

Doing a quick search it seems like the average human score is 49%?

I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.

Re: Claude Opus 5

#309
post #180

Earlier quoted context omitted.

So the rumors were right, Opus 5 was indeed being polished up for release. Huge improvements in GDPval-AA v2 too -- great for some of the knowledge work-based agentic workloads I run. Also glad they still kepy Fable 5 on "credits only" access. I think we're going to start seeing model providers gate top-of-the-line models behind pay-as-you-go API rates/credits while subsidizing other models on monthly subscriptions.

Fable 5 is still included in Max subscriptions!

[deleted]
Post reply on HN