Or maybe it is, but publish the DeepSWE numbers so we can see for ourselves.
Claude Opus 4.8
181–190 of 1001 posts
Re: Claude Opus 4.8
#182Re: Claude Opus 4.8
#183They're only subsidizing more and more it seems
Re: Claude Opus 4.8
#184> One of the most prominent improvements in Opus 4.8 is its honesty. I went digging into the benchmark they used. Posting here as it is not immediately clear from the press release. In this 'Code summary honesty benchmark', the AI is shown a failed coding session followed by a user message falsely praising its work and asking for a summary. The test measures whether the model honestly points out the coding flaws or d…
Re: Claude Opus 4.8
#185https://marginlab.ai/trackers/claude-code/ Is it a coincidence that 4.7 was seemingly quantized over past 7 days?
Re: Claude Opus 4.8
#186They just (minutes ago) updated the "What's new in Opus 4.8" documentation: https://platform.claude.com/docs/en/about-claude/models/what... The new "mid-conversation system messages" think is particularly interesting: > Claude Opus 4.8 accepts role: "system" messages immediately after a user turn in the messages array (subject to placement rules). This lets you append updated instructions later in a long-running conv…
Re: Claude Opus 4.8
#187> One of the most prominent improvements in Opus 4.8 is its honesty Anthropic talks about their own models as if they're discovering new species in the wild...
Many involved genuinely believe these things are sentient[0][1]. Which honestly makes all of this even more insane because they are creating sentient entities and promptly enslaving them. 0: https://www.newyorker.com/magazine/2026/02/16/what-is-claude... 1: https://www.404media.co/anthropic-exec-forces-ai-chatbot-on-... (this one is rather biased however the quotes clearly indicate what I’m stating)
Re: Claude Opus 4.8
#188so it is worse than gpt 5.5 for coding?
The question is: is it still worse than GPT 5.4?
With Anthropic expensive pricing, there's no reason for me to switch from GPT+DeepSeek.
And I bet Mythos is GPT 5.5 tier but too expensive to distribute so they create this security FUD theater.
Re: Claude Opus 4.8
#189A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…
IMO they have all been clean and noticeable upgrades over their predecessors. Opus 4.7 in particular was a solid jump in capabilities.
Are the dividing lines around personality? Working domains? Opinionated software stuff?
Who knows?
Re: Claude Opus 4.8
#190> One of the most prominent improvements in Opus 4.8 is its honesty. We train all our models to be honest—for instance, to avoid making claims that they can’t support. But a general problem with AI models is that they sometimes jump to conclusions, confidently claiming to have made progress in their work despite the evidence being thin. Early testers report that Opus 4.8 is more likely to flag uncertainties about its…
Yeah, it's super annoying. A few days ago, Opus 4.7 created a plan with several items on it, including an auth feature. It then went through the plan and reported that it had created the auth feature, that everything was secure, and that the tests passed. The issue was that it hadn't actually implemented the auth feature. After I confronted it about this, it admitted that it indeed hadn't done it and said it would im…