Live data from Hacker News

Claude Opus 4.7 Model Card

anthropic.com

21–30 of 88 posts

Re: Claude Opus 4.7 Model Card

#21
post #16
post #12

Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.

Both. There's the risk of them instructing a user on how to produce a known formulation (the Anarchist Cookbook solution, as you say), which is irritating but not that problematic. The bigger issue is that they are potentially capable of producing novel formulations capable of producing harm, and guiding someone through this process. That is, consider a world in which someone with malicious desires has access to a mo…

The world has been blessed by two connected things:

1. Smart people have economic opportunities that align them away from being evil

2. People who are evil tend not to be smart.

We're breaking both of these assumptions.

Re: Claude Opus 4.7 Model Card

#22
> The technical error that caused accidental chain-of-thought supervision in some prior models (including Mythos Preview) was also present during the training of Claude Opus 4.7, affecting 7.8% of episodes.

>_>

Re: Claude Opus 4.7 Model Card

#23
post #12

Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.

In the same way that all coding docs are available publicly

Re: Claude Opus 4.7 Model Card

#24
post #16

Earlier quoted context omitted.

Both. There's the risk of them instructing a user on how to produce a known formulation (the Anarchist Cookbook solution, as you say), which is irritating but not that problematic. The bigger issue is that they are potentially capable of producing novel formulations capable of producing harm, and guiding someone through this process. That is, consider a world in which someone with malicious desires has access to a mo…

The world has been blessed by two connected things: 1. Smart people have economic opportunities that align them away from being evil 2. People who are evil tend not to be smart. We're breaking both of these assumptions.

Good. This is how we will force the world to reckon with the isolated, the disgruntled, and "lone wolf" terrorist. Real "sigma males" actually exist, and when they decide "society has to pay" we are all worse off for it. If Ted Kaczynski (quintessential example of a real actual sigma) had been in his prime operating right now, he'd have mail-bombed NeurIPS and ICLR already. I'm not cool with being in crowds of AI professionals right now for physical security reasons given the extreme anti-AI sentiment that exists from nearly everyone outside of the valley: https://jonready.com/blog/posts/everyone-in-seattle-hates-ai...

Re: Claude Opus 4.7 Model Card

#25
So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.

Re: Claude Opus 4.7 Model Card

#26
post #18
post #9

Earlier quoted context omitted.

The Gemma models are at this point. A 31B model that can fit on a consumer card is as good as Sonnet 4.5. I haven't put it through as much on the coding front or tool calling as I have the Claude or GPT models, but for text processing it is on par with the frontier models.

absolutely not on par you're smoking

You make a compelling argument, but thankfully I have data to back up my anecdotal experience

This comparison shows them neck and neck https://benchlm.ai/compare/claude-sonnet-4-5-vs-gemma-4-31b

As Does this one https://llm-stats.com/models/compare/claude-sonnet-4-6-vs-ge...

And the pelican benchmark even shows them pretty close https://simonwillison.net/2026/Apr/2/gemma-4/ https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/

Also this isn't a fringe statement, you can see most people who have done an evaluation agree with me

Re: Claude Opus 4.7 Model Card

#27
post #19
post #18

Earlier quoted context omitted.

absolutely not on par you're smoking

Just to be clear, did you notice the parent said 4.5?

They are also on par in a lot of classification tasks. I did have to actually use gemma4 and fine tune it a bit but that is part of the value add.

Re: Claude Opus 4.7 Model Card

#29
post #16

Earlier quoted context omitted.

Both. There's the risk of them instructing a user on how to produce a known formulation (the Anarchist Cookbook solution, as you say), which is irritating but not that problematic. The bigger issue is that they are potentially capable of producing novel formulations capable of producing harm, and guiding someone through this process. That is, consider a world in which someone with malicious desires has access to a mo…

The world has been blessed by two connected things: 1. Smart people have economic opportunities that align them away from being evil 2. People who are evil tend not to be smart. We're breaking both of these assumptions.

"Smart people have economic opportunities that align them away from being evil"

For some definition of evil, some of the time, ok. But as economic opportunities compound (looking at the behavior of the ultra-rich), it seems there's at least strong correlation in the other direction, if not full-on "root of all evil" causation.

Re: Claude Opus 4.7 Model Card

#30
post #11

Have they effectively communicated what a 20x or 10x Claude subscription actually means? And with Claude 4.7 increasing usage by 1.35x does that mean a 20x plan is now really a 13x plan (no token increase on the subscription) or a 27x plan (more tokens given to compensate for more computer cost) relative to Claude Opus 4.6?

Definitely 13x, at least for now
Post reply on HN