Live data from Hacker News

Claude Opus 5

anthropic.com

201–210 of 1001 posts

Re: Claude Opus 5

#201
post #97

Earlier quoted context omitted.

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!

Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend. Fable is typically used for key planning, architecting, and review tasks. I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.

Eh, not really. Fable does a lot better on coding than Opus 4.8.

Just this past week Fable was able to figure out a couple of small issues for me where Opus was failing to.

Also both are still somewhat bad at UI implementation. Opus more so

Re: Claude Opus 5

#202
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

> This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models

Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying price-war. All one needs to do is raise their price and watch how the competition react.

Such a scheme (and resulting high margins) would be imperilled by the existence of frontier open-weight models in the market, which may be why the reaction to Chinese models may be particularly shrill.

Re: Claude Opus 5

#203
post #41

Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…

Openrouter should ideally kill in this space and make their model agnostic infra like memory, harnesses, chat applications.

Re: Claude Opus 5

#204

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores. I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

If they were benchmaxxing, surely they would score higher than 30% on ARC-AGI.

Re: Claude Opus 5

#205
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

> This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models Gemini also had modest increase before this - don't be surprised when OpenAI also has a "modest increase" with its next release. Cartel-like behaviour doesn't require direct communication when none of the participants are interested in participating in a margin-destroying p…

>Don't be surprised when OpenAI also has a "modest increase" with its next releases; cartel-like behaviour doesn't require overt coordination when none of the participants are interested in participating in a margin-destroying price-war.

No I will be surprised and I'll bet on the fact that prices will keep going down, just like it went ~50% down in the latest GPT 5.6 release.

Re: Claude Opus 5

#206
This stood out to me as a little concerning:

> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.

Re: Claude Opus 5

#209
post #41

Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…

Openrouter should ideally kill in this space and make their model agnostic infra like memory, harnesses, chat applications.

OpenRouter is in acquisition talks with Stripe, fyi

I would expect routers to commodify like tokens.

Re: Claude Opus 5

#210
post #6

> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5 Ok then so what's the point?

You are being downvoted for a fair question and others are extremely wrong and confident.

The point is that Opus 5 is the best they can do without needing classifiers and absurdly broad safeguards.

Post reply on HN