Live data from Hacker News

Claude Opus 5

anthropic.com

331–340 of 1001 posts

Re: Claude Opus 5

#331
post #305
post #90

Earlier quoted context omitted.

I guess character? Fable is more friendly and curious while opus is a bit more deliberate and conservative.

Surely that's tunable? OpenAI lets you tune response characteristics.

Yes, but there is a baseline.

Re: Claude Opus 5

#332
Doing testing with it now, specifically for image->html conversion.

Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).

Opus' results seem to be more accurate than Fable, following the design source of truth better.

Example results:

Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...

Opus 5 build: https://html.non.io/solaraOpus/

Fable 5 build: https://html.non.io/solara/

Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).

Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.

Re: Claude Opus 5

#333
post #307

But why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed. I've never trusted on model cards though. I'm sorry.

Another benefit is that fast mode can be used on subscription, but Anthropic's won't

Exactly! And they should also release the new inference engine in this month. Anyway, I am curious to try Opus 5, considering that previous versions (e.g., 4.8) were disappointing

Re: Claude Opus 5

#334
post #315

Earlier quoted context omitted.

what's the threshold for model routing where you're willing to trust the router? For coding my own work I don't trust the model router, and it would have to be shown to be to save a real dollar amount. From a buying perspective it's a hard sell to save x but lose out on bugs you are probably introducing at an unquantifiable severity and frequency. How much is it worth to hedge your bets by doing every single inferenc…

you trust the service provider but not the router? weird, but ok

I don't want to save money so badly that I'd possibly undercut the quality of the code that gets created.

*edit to add: that code quality (or lack of quality) is it's own cost

Re: Claude Opus 5

#335
The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

Re: Claude Opus 5

#336
post #164

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

It seems they are trying to thread a needle here - they want to say it's very strong, but apparently this time do not want to invite extra government scrutiny.

They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.

"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."

Re: Claude Opus 5

#337

Looking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

With Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).

Re: Claude Opus 5

#338

The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

Opus is better than Sonnet -- an Opus is longer than a Sonnet

Re: Claude Opus 5

#339
post #236
post #41

Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…

Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable. Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability

Model routing by the model itself requires the model to pull in a lot of context and it's likely more efficiently just done by people with the context already in their head, even assuming the model is perfect at routing (which last I checked, Claude definitely isn't). I wouldn't trust ML model routers.

Re: Claude Opus 5

#340
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

I applied to their reliability team but never heard back. I would love to help them solve this problem!
Post reply on HN