Claude Fable 5.1 and Claude Mythos 5.1
731–740 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#732Earlier quoted context omitted.
Now that it's a solved benchmark, can we get the animated version?
I didn't want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high): llm logs -cx | llm -m claude-fable-5.1 -s 'animate this' Here's the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... It's excellent!
Re: Claude Fable 5.1 and Claude Mythos 5.1
#733Earlier quoted context omitted.
From Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more. Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price. https://artificialanalysis.ai/mod…
From my limited testing of just 2 hours, reasoning output of 5.1-max is at least 7x of 5-max, on the same project and comparable prompts. It reasoned for ~2 minutes trying to figure out an appropriate directory name. I've never seen 5-max do that. Could be a misconfiguration though.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#734Re: Claude Fable 5.1 and Claude Mythos 5.1
#735Earlier quoted context omitted.
I sometimes have that feeling too, then ask another LLM to do a code and vulnerability review and OMG: rookie mistakes, over complications and security gaps even a 1st year student would not make regularly. So.. one more year of untreated bipolar AI psychosis I guess..
I think these kinds of comments really need to say which LLM that is. There's an enormous difference in skill between the frontier ones and say the Google search AI.
They all work, they all are "good", they all are both "smart" and commit incredible basic mistakes a fair amount of times.
Then there's the cost situation..
Re: Claude Fable 5.1 and Claude Mythos 5.1
#736Re: Claude Fable 5.1 and Claude Mythos 5.1
#737When paying by the token, don't the labs have a strong incentive to make the model as verbose as possible?
Re: Claude Fable 5.1 and Claude Mythos 5.1
#738Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…
That wasn’t Anthropic. Clearly not a well informed take.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#739(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
Does the new writing style now have EU level watermarks?
Re: Claude Fable 5.1 and Claude Mythos 5.1
#740Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…