Live data from Hacker News

Claude Opus 5

anthropic.com

121–130 of 1001 posts

Re: Claude Opus 5

#121
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

The final user-facing responses are usually a tiny fraction of the total tokens used over the course of a given conversation turn. When you're doing any real work, reasoning and tool uses constitute the overwhelming majority of the tokens in / out... not the final user-facing response.

Re: Claude Opus 5

#122

Judging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now. Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon. AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported…

I think at some point we might see something akin to LTS releases, especially if/when capability improvement slows to a crawl.

Re: Claude Opus 5

#123

The wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions. > we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities I wonder if Anthropic would still intentionally nerf their…

Opus 4.8 was intentionally nerfed and that was before the government took action against Fable

Re: Claude Opus 5

#124

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete. --------------- Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Op…

I think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.

Re: Claude Opus 5

#126
post #6

> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5 Ok then so what's the point?

Fable 5 is NOT included in Claude Pro subscription

Aren’t they planning to remove it even from max and keep it only credit based? OpenAI will be happy if that would happen

Re: Claude Opus 5

#128
post #97
post #17

https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!

Using the most expensive model for all of your agentic coding work hasn’t been good practice for a long time. Not unless you have infinite money to spend.

Fable is typically used for key planning, architecting, and review tasks.

I think this is a case where you don’t understand the use case, not that the marketing department is making mistakes.

Re: Claude Opus 5

#129
post #85
post #48

Earlier quoted context omitted.

Because their model previously got blocked by the government for this and they don't want a repeat?

5.6-Sol is a lot more permissive than Opus/Fable even w/ CVP (once you sign your soul away to Palantir via Persona, anyway), while maintaining better capabilities

5.6-sol in a single prompt was able to discover a zero-day in a web application (with no sourcecode provided, only known api urls) and I do not even have /cyber verification on my general purpose account. I wasn't even really tryign to "find" a zero-day it was just looking for bypassing a restriction... Instead of spending 10 minutes filling in a form I ended up having to spend an hour drafting a report and sending an email.

Re: Claude Opus 5

#130

That's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?

Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

Post reply on HN