Live data from Hacker News

Claude Opus 5

anthropic.com

531–540 of 1001 posts

Re: Claude Opus 5

#531

This is cool, but I wish we could stop building landing pages to assess the intelligence of these models. There is much more to them than that. There are infinite number of complicated things that require a crap-ton of intelligence (biological or digital). The most fascinating of these for me these days is large scale migrations. Projects that are so ginormous and risky that many teams have either given up on them, o…

Outsourcing blame. This is the killer app of AI

Re: Claude Opus 5

#532
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

This is part of why I switched to Grok 4.5 I don't need more powerful models, I need one that responds fast enough that my attention doesn't wander to other tasks. Grok 4.5 is so fast I can just use it in-band without swapping to other tasks. Slower than Opus 4.8, which was already miserably slow, is indeed a step in the wrong direction.

After actually using for most of a day now, it really seems not just slightly slower than 4.8, but way slower. Even relatively easy code refactoring tasks seem to take a while.

Re: Claude Opus 5

#533

What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.

I actually do read them. Not in severe detail, but not casually either. 150 pages is really not very long and there doesn't seem to be too much bloat. (I would cut out the moral personhood stuff but that's a political/ideological thing). This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.

You mean moral patienthood?

There are many ways to read something, model cards are usually skimmed.

Re: Claude Opus 5

#534
post #228
post #155

Earlier quoted context omitted.

Also the cost per task. It appears to be significantly cheaper, cheaper than sonnet!

The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3. https://artificialanalysis.ai/?cost=cost-per-task

[deleted]

Re: Claude Opus 5

#535

Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.

I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going

Or another way to see it is that current models are AGI as it was defined before, and the goal post is being moved.

Re: Claude Opus 5

#536
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

In my tests, it averages to much cheaper than Opus 4.8 on real tasks on account of being smarter and more token efficient.

I have a benchmark to build a game engine from a set of written instructions. It's a little tricky. Opus 4.8 did it in 470k tokens at a cost of $1.29 vs Opus 5 in 179k tokens for $0.33. (Fable 5 did it in 245k for $0.95)

Though if you really want to cut costs, Tencent's Hy3 model also got it right and did it in 283k tokens for $0.03

Re: Claude Opus 5

#537
post #524

Earlier quoted context omitted.

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

That's a plausible explanation but I'm not seeing evidence for it. I have a personal benchmark suite of 14 real, non-public tasks. Opus 5 and Fable tied on 10, Opus won on 3, and Fable won on 1. It's a really strong model.

This is an interesting idea, without giving away your benchmarks specifically what kind of stuff do you test? I might try to assemble something myself, it's so hard to determine model quality from the system card these days.

Re: Claude Opus 5

#538
Alongside this release I seem to have lost all thinking traces from all models - now it only generates a one-line summary similar to Gemini. I'm guessing this is an anti distillation measure? I'm surprised to see no one else complaining about this, it's a significant reduction in usefulness not being able to explore alternative angles that the model discarded in the final output.

Re: Claude Opus 5

#539
post #535

Earlier quoted context omitted.

I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going

Or another way to see it is that current models are AGI as it was defined before, and the goal post is being moved.

they are definitely not agi as it was ever defined. they’re only a bit more capable than they were a year ago. they crossed over from interesting crap to useful tool recently but really only for software

Re: Claude Opus 5

#540

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

Have you played Arc 3? It seems like more of a simple optimization problem (think Sokoban) than anything approaching fluid intelligence. Whether a multi hundred billion dollar company would spend time benchmaxxing a highly publicized benchmark that claims to confer AGI is an exercise left to the reader, but I doubt Claude Plays Pokemon is suddenly going to get past Mt. Doom now.

What? Claude plays Pokemon one-shots the entire game without any harness other than claude code and game screenshots
Post reply on HN