Who cares about Sonnet? I want to know about Fable. Are the export restrictions really going to be permanent?
It’s supposed to happen when Anthropic introduces identification, which I believe is planned for mid-July.
Claude Sonnet 5
451–460 of 822 posts
Re: Claude Sonnet 5
#452> Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. And Opus 4.8 is still cheaper for a higher pass rate (much less open weight models like GLM 5.2) so not sure why I'd use Sonnet except on the…
"Lower ability to perform cybersecurity-related tasks" makes me super concerned it will leave my codebase like Swiss cheese for any American granny with access to Fable 5, when we non-American Brits, or rest-of-worlders, don't have access to it to clean our codebases.
Re: Claude Sonnet 5
#453> Claude Opus 4.7 and later Opus models, Claude Fable 5, Claude Mythos 5, Claude Mythos Preview, and Claude Sonnet 5 use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. Claude Sonnet 4.6 and earlier models use the previous tokenizer.
Re: Claude Sonnet 5
#454Wonder if the whole cyber paranoia leads to their models ultimately generating less secure code. After all, if it has the ability to generate safe code, it would imply that it knows something about cybersecurity, which could surely be used to hack all the banks in the world.
Re: Claude Sonnet 5
#455I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like tests [0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
Re: Claude Sonnet 5
#456In my early tests tonight, Sonnet 5 is a LOT better out of the box. It's one-shotting complex instructions. It also recovered independently from bad instructions that led to an uninformative 400 error by using its schema-fetching tool to figure out there were was too much input.
If I have to gripe about something: it interpreted another impossible instruction by quietly discarding the input in question. But, the way it did it is... kinda exactly what anybody else would do, if they weren't in a position to change the implementation.
This is, obviously, early days but I'm impressed.
Re: Claude Sonnet 5
#457I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like tests [0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
Still one of the most intelligent models overall, most likely to get any question you ask correctly (without tools).
Re: Claude Sonnet 5
#458GLM 5.2 is better and cheaper. Maybe they are trying to embarrass Trump by making it look like we are losing to China.
Re: Claude Sonnet 5
#459I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45/100, sometimes makes invalid tool calls - Puzzle Solving — score 77, flubs carwash-like tests [0]: https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
Re: Claude Sonnet 5
#460Earlier quoted context omitted.
I always use Opus 4.8 at max effort for everything. The $20 subscription didn't have enough tokens, but the $100 one had too many of them. So now I just max out Opus in order to maintain 100% weekly utilization.
Agreed I think your strategy is optimal. This is what I landed on as well