Live data from Hacker News

Claude Opus 5

anthropic.com

411–420 of 1001 posts

Re: Claude Opus 5

#411
post #40

I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.

Fable 5 is assumed to be a larger model. It seems plausible to me that RL improvements allowed Anthropic to improve on Opus 4.8, similar to how OpenAI substantially improved upon GPT 5.5 with 5.6 Sol. Fable 5.1 and GPT-6 are rumored to launch in August, presumably bringing those improvements to the larger models.

An Anthropic "leak" back in March said that "'Capybara' is a new name for a new tier of model: larger and more intelligent than our Opus models — which were, until now, our most powerful". A second version of the leak had it referring to Claude Mythos rather than Capybara.

I don't know how systematic Anthropic are about their versioning - I'd have guessed that major version number increases (4.x -> 5.x) reflect different base models (different pre-training runs), in which case Opus 5 would be a distilled version of the Fable 5 base model (but without the cyber exploit post-training), rather than being Opus 4.8 with additional post-training, but who knows? I don't believe Anthropic have said anything about this.

Re: Claude Opus 5

#412

Earlier quoted context omitted.

Opus is better than Sonnet -- an Opus is longer than a Sonnet

> an Opus is longer than a Sonnet And people know this? I didn't. I am not into music or poetry so these are not terms I am familiar with.

I guess you are one of today's lucky 10,000 [0]

[0] https://xkcd.com/1053/

Re: Claude Opus 5

#413

Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.

So unless doomsday actually happens then you're unhappy with the warning - is that right? You see false promises of apocalypse as marketing?

It’s either advertising, or they’re idiots, because the apocalypse keeps not happening. Either way, it’s not worth listening to them.

Re: Claude Opus 5

#414
post #51

Better than Fable 5 on all but 3 evals. Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?

Pretty sure Mythos and Fable have way more params, but they've just been able to use the synthetic data off of them to get the leap in quality from Opus.

So, not a distilled version of Mythos or Fable, but those models likely helped a lot in the post training phase of Opus.

Re: Claude Opus 5

#415
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

Is it some Claude Team/Enterprise only problematic? I'm using two 20x Max accounts almost non-stop (Fable/Opus) for 1.5 years at this point, zero issues with both client and infra sides (from US and in travels). When I'm reading such messages it feels like either I'm lucky or it's a part of some campaign.

I’m on the biggest max plan. It is riddled with annoying bugs for me, only been using it for a little over a month. Settings screen flashes randomly. But most annoyingly: sometimes when forking chats or sometimes for no reason, the UI just straight up eats my previous messages. The model is still aware of them and can recount them if I ask but the visible history is gone. And that’s not even all of them. Fable 5 is just too good that I put up with it but it seriously raises concerns for me that even with infinite compute these companies can’t even deliver a functional chat UI.

Re: Claude Opus 5

#416
The chaos appears to be tamed for now.

From the system card [1]:

  The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code.
  If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing.
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...

Re: Claude Opus 5

#417
post #382

Earlier quoted context omitted.

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

God damn, we are living in the future. I love this so much. Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again. This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is). This is fun and it's go…

No we are not living in future. Design is ugly, and immediate put off because it smells AI.

Re: Claude Opus 5

#418

Opus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2]. Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason? [0]: https://artificialanalysis.ai/#intelligence [1]: https://platform.claude.com/docs/en/about-claude/pricing…

Benchmarks have gotten great, but they're still a proxy for the real world. The 3 GPT 5.6 models are also further apart in reality than the numbers suggest. That said, I'm still mighty impressed how good Luna is for the price. Highly underrated model.

I have been trying to build something that captures the behavioral element of different models, but it's kinda tough.

Re: Claude Opus 5

#419
post #155

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

I can't believe they released the charts they did.

It basically shows that Sol absolutely demolishes Fable at every part of the cost curve for coding for the same level of quality.

Opus is competitive. It just has a higher level of quality / higher cost to start.

Re: Claude Opus 5

#420

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

insane pricing: " Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8)"

I think for the value of the outputs that’s still a good deal. Keeping the same price as the prior model makes sense to me. That is if the model size is about the same in the cost to serve has not substantially changed. Now I would have expected efficiency gains for inference, but there is no way to know as a customer.

At the end of the day, they have established a strong brand and if they can get away with a 95%+ gross margin on inference entirely from the status premium, then I suppose that’s good for them. Apple does the same thing, and I don’t fault them for it.

Post reply on HN