Live data from Hacker News

Claude Opus 5

anthropic.com

371–380 of 1001 posts

Re: Claude Opus 5

#371

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.

Re: Claude Opus 5

#373

The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

Fable is better than Opus which is better than Sonnet which is better than Haiku. They’re basically just sizes. Though it gets even more confusing because they also have effort levels so it’s not really possible to call one fast and one slow since Fable on Medium will be faster than Opus on Max. I agree it’s confusing, and now OpenAI is following Anthropic’s lead with their new naming (Sol, Terra, Luna).

But according to their benchmarks Opus 5 outscores Fable 5 on basically everything. So which one is “better”?

Re: Claude Opus 5

#374
post #97

Earlier quoted context omitted.

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!

Best can describe multiple things. Almost as good for half the cost is something I'm very comfortable describing that way.

> Almost as good for half the cost is something I'm very comfortable describing that way.

It's also not unusual in this context - many people describe the Chinese models as "best", because it's 80% as good for 20% of the price (or similar).

Re: Claude Opus 5

#375

Is Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?

No, they're different models. Knowledge cutoff has been updated.

Re: Claude Opus 5

#376
It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?

Re: Claude Opus 5

#377
post #296
post #236

Earlier quoted context omitted.

Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable. Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability

This doesn’t seem obviously true, eg an Anthropic model will never route to Kimi even if it were best suited for a particular task.

Why should it? An Anthropic model is architecturally optimized for Anthropic models, routing it to Kimi makes zero sense

Re: Claude Opus 5

#379

Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.

So unless doomsday actually happens then you're unhappy with the warning - is that right? You see false promises of apocalypse as marketing?

My point is why the sudden change in tone? I’m not dismissing the models’ capabilities.
Post reply on HN