Live data from Hacker News

Claude Opus 5

anthropic.com

551–560 of 1001 posts

Re: Claude Opus 5

#551
post #446

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

Yes - THIS! I can't even believe how exhausting it is to read. I'm not sure why or what changed in Fable. Did they do this writing-style output to give it more token compression during/for training or to prefer output for less money?

I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.

I'm guessing it's just not enough time doing RL on human feedback.

Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle

"Early in RL verbose, grammatical" (if you search) :

We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².

vs. Post RL

We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …

I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.

Re: Claude Opus 5

#552
> An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.

This is a pretty common trading firm internship project funnily enough.

Re: Claude Opus 5

#553
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

Pure speculation, but I've noticed drawbacks to the models on high effort. I interact mostly through prompts rather than agents so I sometimes see where their reasoning falls short. A model on high effort has longer output and can get hyperfocused on irrelevant details, maybe increasing the surface area for mistakes. I haven't used other effort levels extensively yet but I've supposed that medium may have more balance.

Re: Claude Opus 5

#554

Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?

RL. Lots of RL

Yes. I’m 99% sure arc agi 3 will be saturated like 1 and 2. In less than a year. And they will come up with one more.

Re: Claude Opus 5

#555
post #452

Earlier quoted context omitted.

And to the guardrails of Fable: https://x.com/cheatyyyy/status/2080693704290140330

That tweet says: > Opus 5 can silently fallback to Opus 4.8 (without any notice) on the serverside if you hit a guardrail But https://support.claude.com/en/articles/16049681-why-claude-s... says (emphasis mine): > These checks cause Claude to _visibly_ fallback from Opus 5 to Opus 4.8 [...] You'll see a notice explaining that the model switched, and the response will be labeled with the model that answered. So who is…

It's showing you're switched to 4.8, i just hit that while doing security research.

Re: Claude Opus 5

#556
post #332

Doing testing with it now, specifically for image->html conversion. Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models). Opus' results seem to be more accurate than Fable, following the design source of truth better. Example results: Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we... Opus 5 build: https://html.non…

I was curious to see how open weight models would do on this task so I passed in a screenshot of your source of truth and here's what 2 of the best code-generation models that allow image inputs do:

Inkling (not too great): https://cdn-uploads.huggingface.co/production/uploads/608b8b...

Kimi 2.7 (really well, esp. note that this is the predecessor model, not the latest Kimi3): https://cdn-uploads.huggingface.co/production/uploads/608b8b...

Here's how I tested them: https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...

https://huggingface.co/spaces/abidlabs/vlm-screenshot-to-web...

Re: Claude Opus 5

#557
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

It's "Cost per task", so perhaps it burns tokens too quickly, trying to "do a better job". Over-engineering :)

Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.

Completely subjective, of course.

Re: Claude Opus 5

#558

Looking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...

AA isn't the best way to measure relative cost in real world use because some of those benchmark questions are extremely hard for the models. Some models give up quickly on hard questions, other models spin their wheels for a long time before declaring defeat (or getting the answer on token 200k!).

A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.

Re: Claude Opus 5

#559

Earlier quoted context omitted.

With Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).

You can opt out of training. If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.

All providers are equally trustworthy :)

Re: Claude Opus 5

#560

The naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.

You're confused by this? Fable > Opus > Sonnet > Haiku. Sort by price. Most expensive = better. Are you an alien? lol
Post reply on HN