Claude Opus 4.8
51–60 of 1001 posts
Re: Claude Opus 4.8
#52> One of the most prominent improvements in Opus 4.8 is its honesty. We train all our models to be honest—for instance, to avoid making claims that they can’t support. But a general problem with AI models is that they sometimes jump to conclusions, confidently claiming to have made progress in their work despite the evidence being thin. Early testers report that Opus 4.8 is more likely to flag uncertainties about its…
Re: Claude Opus 4.8
#53Not half bad!
Re: Claude Opus 4.8
#54So GPT 5.6 tomorrow, then?
Re: Claude Opus 4.8
#55> One of the most prominent improvements in Opus 4.8 is its honesty Anthropic talks about their own models as if they're discovering new species in the wild...
Re: Claude Opus 4.8
#56> One of the most prominent improvements in Opus 4.8 is its honesty. We train all our models to be honest—for instance, to avoid making claims that they can’t support. But a general problem with AI models is that they sometimes jump to conclusions, confidently claiming to have made progress in their work despite the evidence being thin. Early testers report that Opus 4.8 is more likely to flag uncertainties about its…
Don't play to the sci-fi "this thing's trying to outsmart me" tropes.
Re: Claude Opus 4.8
#57Anthropic has now upgraded their Claude slot machine to version 4.8. Time to gamble even more tokens at the Anthropic casino.
> Claude can plan the work and then run hundreds of parallel subagents in a single session (and with Opus 4.8, the agents can run for even longer).
Re: Claude Opus 4.8
#58In our work we asked several frontier AIs to come up with an API we needed. We compared Opus 4.7 and GPT-5.5 (among others). Opus 4.7 came up with the most creative and intelligent API design that pleasantly surprised us, especially given that GPT-5.5 was passing it on various coding benchmarks.
What I noticed is that we don't have a commons benchmark to measure "creativity" and "ingenuity", and in some ways such a benchmark would conflict with the common IFBench benchmark. Yet this is a very important skill when designing systems. I'm glad to see Anthropic putting thought into it, and would love to see a public benchmark for this that other models could compare themselves to.
[1] https://cdn.sanity.io/files/4zrzovbb/website/c886650a2e96fc0...
Re: Claude Opus 4.8
#59Crazy they bring up honest, when Claude models are literally known for straight up lying about things it has done and tries to act like it did what you asked.
Less than other frontier models. Which is scary honestly.
You tell it too research a repo to find a piece of code it will. Claude will just read the README and guess.