Live data from Hacker News

Claude Opus 4.7 Model Card

anthropic.com

71–80 of 88 posts

Re: Claude Opus 4.7 Model Card

#72

This is an interesting document, in that it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card. I surmise that someone at the top put the Mythos release on hold, and the product team was told "ship this other interim step model instead. quickly." I wonder if 4.7 will be seen as a net step-up in quality; there are some regressions noted in the document, and it's clearly substant…

Yeah, the section expanding on how they evaluated Mythos internally is a bit baffling considering how irrelevant it is.

Re: Claude Opus 4.7 Model Card

#73
post #38

So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.

To be honest, I think it's just a more honest score of what Opus 4.6 actually was. Once contexts get sufficiently large, Opus develops pretty bad short term memory loss.

You can support very long context windows if you don’t mind abysmal recall rate.

Re: Claude Opus 4.7 Model Card

#74
post #33

Earlier quoted context omitted.

Could this be because they've found the 1m context uneconomical (ie costs too much to serve, or burns through users quota too quickly causing complaints), and so they're no longer targeting it as a goal

Opus 4.7 is also worse at 256K context. Go look at page 195 and page 196. It is across the board regression, not just 1M context.

Thanks, interesting. Does this make it more surprising that the other benchmarks have improved? I'm not sure I understand the benchmarks well enough - but I'm wondering whether with agentic workflows it's possible to get away with a smaller more focussed context (and hence lower cost) whilst achieving the same or better performance, because of agentic model's ability to decide what the put in context as they work

Re: Claude Opus 4.7 Model Card

#75
post #33

Earlier quoted context omitted.

Could this be because they've found the 1m context uneconomical (ie costs too much to serve, or burns through users quota too quickly causing complaints), and so they're no longer targeting it as a goal

Opus 4.7 is also worse at 256K context. Go look at page 195 and page 196. It is across the board regression, not just 1M context.

what's all this mean in real world use?

Re: Claude Opus 4.7 Model Card

#76
post #26
post #18

Earlier quoted context omitted.

absolutely not on par you're smoking

You make a compelling argument, but thankfully I have data to back up my anecdotal experience This comparison shows them neck and neck https://benchlm.ai/compare/claude-sonnet-4-5-vs-gemma-4-31b As Does this one https://llm-stats.com/models/compare/claude-sonnet-4-6-vs-ge... And the pelican benchmark even shows them pretty close https://simonwillison.net/2026/Apr/2/gemma-4/ https://simonwillison.net/2025/Sep/29/claud…

if you look at the details of the numbers of the benchmarks that you shared, Sonnet 4.5 crushes gemma 4. Somehow the first link doesn't run Sonnet on the multi modal benchmark, that's why the top score looks close, it beats Gemma at every benchmark they actually ran. The arena in the second shows that it actually destroys Gemma 4 as well, not close

Re: Claude Opus 4.7 Model Card

#78
post #16

Earlier quoted context omitted.

Both. There's the risk of them instructing a user on how to produce a known formulation (the Anarchist Cookbook solution, as you say), which is irritating but not that problematic. The bigger issue is that they are potentially capable of producing novel formulations capable of producing harm, and guiding someone through this process. That is, consider a world in which someone with malicious desires has access to a mo…

The world has been blessed by two connected things: 1. Smart people have economic opportunities that align them away from being evil 2. People who are evil tend not to be smart. We're breaking both of these assumptions.

> 1. Smart people have economic opportunities that align them away from being evil

for now

Re: Claude Opus 4.7 Model Card

#79

Model Welfare? Are they serious about this? Or is it just more hype? I really don't trust anything this company says anymore. "We have a model that is too dangerous to release" is like me saying that I have a billion dollars in gold that nobody is allowed to see but I expect to be able to borrow against it.

Maybe referring to it as welfare is odd, but these points are important. It isn't a good look to have a model that tends to get into self-deprecating loops like one of Google's older models, it's an even worse look and potential legal liability if your model becomes associated with a suicide. An overly negative chat model would also just be unpleasant to use.

With the weights being mostly opaque, these kinds of evaluations are an important piece of reducing the harm an AI model can cause.

Re: Claude Opus 4.7 Model Card

#80
post #11

Have they effectively communicated what a 20x or 10x Claude subscription actually means? And with Claude 4.7 increasing usage by 1.35x does that mean a 20x plan is now really a 13x plan (no token increase on the subscription) or a 27x plan (more tokens given to compensate for more computer cost) relative to Claude Opus 4.6?

They have communicated it as 5x is 5 x Pro, and 20x is 20 x Pro (I haven’t looked lately so not sure if that’s changed). They have also repeatedly communicated that the base unit (Pro allotment) is subject to change and does change often. As far as I can tell, that implies there is no guarantee that those subscriptions get some specific number of tokens per unit of time. It’s not a claim they make.

I think as far as the maybe more important weekly allotment Max 5 is 10x Pro and Max 20 is 20x Pro. For the 5 hour window it is as the names would suggest though.
Post reply on HN