So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.
Claude Opus 4.7 Model Card
31–40 of 88 posts
Re: Claude Opus 4.7 Model Card
#32Haiku not getting an update is becoming telling. I suspect we are reaching a point where the low end models are cannibalizing high end and that isn't going to stop. How will these companies make money in a few years when even the smallest models are amazing?
It seems to be a rule that older models are more expensive than newer ones. The low end models have higher $CPT and worse output. I wonder if the move is to just have one model and quantize if you hit compute constraints
It isn't. Gemini has gotten more expensive with each release. Anthropic has stayed pretty similar over time, no? When is the last time OpenAI dropped API prices? OpenAI started very high because they were the first, so there was a ton of low hanging fruit and there was much room to drop.
Re: Claude Opus 4.7 Model Card
#33So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.
Agreed, I appreciate the transparency (and Anthropic isn't normally very transparent). It's also great to know because I will change how I approach long contexts knowing it struggles more with them.
Re: Claude Opus 4.7 Model Card
#34So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.
Re: Claude Opus 4.7 Model Card
#35Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.
On top of LLMs reducing the cost/difficulty, the other reason biological and chemical weapons are such a worry is their asymmetric character — they are much much easier and cheaper to produce and deploy than they are to defend against.
Re: Claude Opus 4.7 Model Card
#36Earlier quoted context omitted.
The world has been blessed by two connected things: 1. Smart people have economic opportunities that align them away from being evil 2. People who are evil tend not to be smart. We're breaking both of these assumptions.
"Smart people have economic opportunities that align them away from being evil" For some definition of evil, some of the time, ok. But as economic opportunities compound (looking at the behavior of the ultra-rich), it seems there's at least strong correlation in the other direction, if not full-on "root of all evil" causation.
So much infrastructure is very soft because the evil people aren’t smart enough to conceive of or conduct an attack.
Re: Claude Opus 4.7 Model Card
#37Re: Claude Opus 4.7 Model Card
#38So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.
Re: Claude Opus 4.7 Model Card
#39Re: Claude Opus 4.7 Model Card
#40This reads more like an advertisement for Mythos, on the first glance