Live data from Hacker News

Claude Opus 4.7 Model Card

anthropic.com

61–70 of 88 posts

Re: Claude Opus 4.7 Model Card

#61
This is an interesting document, in that it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card.

I surmise that someone at the top put the Mythos release on hold, and the product team was told "ship this other interim step model instead. quickly."

I wonder if 4.7 will be seen as a net step-up in quality; there are some regressions noted in the document, and it's clearly substantially worse than Mythos, at least according to its own model card. Should be an interesting few months -- if I were at oAI I'd be rushing to get something out that's clearly better, and pressing for weakness here.

Re: Claude Opus 4.7 Model Card

#62

This is an interesting document, in that it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card. I surmise that someone at the top put the Mythos release on hold, and the product team was told "ship this other interim step model instead. quickly." I wonder if 4.7 will be seen as a net step-up in quality; there are some regressions noted in the document, and it's clearly substant…

What makes you think that? "it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card"

Re: Claude Opus 4.7 Model Card

#63

So Opus 4.7 is measurably worse at long-context retrieval compared to Opus 4.6. Opus 4.6 scores 91.9% and Opus 4.7 scores 59.2%. At least they're transparent about the model degradation. They traded long-context retrieval for better software engineering and math scores.

No: https://x.com/bcherny/status/2044821690920980626

Re: Claude Opus 4.7 Model Card

#64
The model card doesn't mention if this revision will continue to make up and fan vicious conspiracy theories like the prior one does.

I've getting a small but steady stream of harassment from mentally ill people who get spun up on crazy conspiracy theories and claude is all too willing to tell them they are ABSOLUTELY RIGHT, encourage them to TAKE ACTION, and telling them that people who disagree are IN ON IT.

The other major AI LLM services will shut down the deflect to be less crazy or shut down conversation entirely, -- but it seems claude doesn't. Anthropic is probably the worst about prattling on about safety but it seems like their concern is mostly centered on insane movie plot threats and less concerned about things with more potential for real harm.

I've complained to anthropic with no response.

Re: Claude Opus 4.7 Model Card

#65
post #62

This is an interesting document, in that it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card. I surmise that someone at the top put the Mythos release on hold, and the product team was told "ship this other interim step model instead. quickly." I wonder if 4.7 will be seen as a net step-up in quality; there are some regressions noted in the document, and it's clearly substant…

What makes you think that? "it reads like a Claude Mythos model card that was hastily edited to be an Opus 4.7 model card"

There are more mentions of Mythos than 4.6. Mythos results are nearly everywhere, and vastly exceed 4.7's capacity in almost every case. There are sections that report only research on Mythos, none on 4.7. E.g. user surveys about how beneficial Mythos is internally at Anthropic.

Re: Claude Opus 4.7 Model Card

#67
post #32
post #10

Earlier quoted context omitted.

It seems to be a rule that older models are more expensive than newer ones. The low end models have higher $CPT and worse output. I wonder if the move is to just have one model and quantize if you hit compute constraints

> It seems to be a rule that older models are more expensive than newer ones. It isn't. Gemini has gotten more expensive with each release. Anthropic has stayed pretty similar over time, no? When is the last time OpenAI dropped API prices? OpenAI started very high because they were the first, so there was a ton of low hanging fruit and there was much room to drop.

I'm talking about gross margins, not revenue.

It's well known that GPT-4 is much more expensive to operate than the GPT-5 family.

Of course they won't drop the prices; it's pure profit if they make models more efficient.

Re: Claude Opus 4.7 Model Card

#68
Can someone please explain the point of these incremental upgrades? Just release one model. Then maybe do a .5. Then do the next version.

What is the justification for .4.5.6.7.8.9 when the difference isn't measurable and it destroys productivity because they test the next increment on the previous one without customer consent?

Re: Claude Opus 4.7 Model Card

#69
post #12

Dumb question but why are chemical weapons always addressed as a risk with llms? Is the idea that they contain how to make chemical weapons or that they would guide someone on how? Would there not already be websites that contain that information? How is an llm different, i guess, from some sort of anarchist cookbook thing.

Probably also a bit of liability. After all its been trained on a dataset that includes a long running joke of trying to trick people on the internet to unknowingly create chlorine gas.

Re: Claude Opus 4.7 Model Card

#70
post #4

Haiku not getting an update is becoming telling. I suspect we are reaching a point where the low end models are cannibalizing high end and that isn't going to stop. How will these companies make money in a few years when even the smallest models are amazing?

Google is putting a lot of research into small models. Most of my AI budget is now going to small models because I am doing lots of tiny tasks that the small models do great with. I would think a decent chunk of Goog's API revenue probably comes from their small models.
Post reply on HN