Live data from Hacker News

Claude Opus 5

anthropic.com

471–480 of 1001 posts

Re: Claude Opus 5

#471

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

[deleted]

Re: Claude Opus 5

#472
post #155

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

Also the cost per task.

https://www.vals.ai/benchmarks/vals_index

!!! Vals !!!

Vals Index Opus 4.8 > 5.0 goes from $2.90 to $8.54, for 4% gain ... That is a massive cost increase. Sure, 20% cheaper then Fable, but that is a 3x price increase compared to Opus 4.8 in that test.

https://artificialanalysis.ai/models/claude-opus-5 https://artificialanalysis.ai/models/claude-opus-5#price-cos...

!!! artificial analysis !!

Cost per task is second highest, right below Fable.

* Fable: $2.75

* Opus 5.0: $2.03

* Opus 4.8: $1.80

* GPT 5.6 Sol: $1.04

* Kimi K3: $0.95

Looks like interest levels of cherry picked cost in their report. Cheaper model, clearly NOT. More expensive in both benchmarks.

Re: Claude Opus 5

#473

Looking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

As always, it requires evaluation with your work because I’m often finding grok to be much more expensive than the price would lead you to believe.

There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.

Re: Claude Opus 5

#474
post #97

Earlier quoted context omitted.

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!

At half the price and less likely to auto-downgrade, it sounds like a reasonable claim.

> At half the price and less likely to auto-downgrade, it sounds like a reasonable claim

Two benchmarks (artificial analysis and vals) show a increase in cost (a insane increase for vals compared to Opus 4.8).

Already posted this before, so here is the link.

https://news.ycombinator.com/item?id=49041158

Re: Claude Opus 5

#476

Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.

I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going

Yes 18 months ago it seemed like AGI was being promised every other week, and now I don't see any of those headlines.

Re: Claude Opus 5

#477
post #273

For anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4 . And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow

Cinematic video link is incorrect. Podcast link is correct.

Re: Claude Opus 5

#479
post #228
post #155

Earlier quoted context omitted.

Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3. https://artificialanalysis.ai/?cost=cost-per-task

Testing at max effort likely doesn't produce optimal results.

Re: Claude Opus 5

#480
post #452

I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0]. > "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1] On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2]. 0: h…

And to the guardrails of Fable: https://x.com/cheatyyyy/status/2080693704290140330

That tweet says:

> Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail

But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine):

> These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered.

So who is right? I know for Fable I am visibly told, is this tweet trying to say it is silent against what Anthropic is saying?

Post reply on HN