Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

731–740 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#732
post #304
post #248

Earlier quoted context omitted.

Now that it's a solved benchmark, can we get the animated version?

I didn't want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high): llm logs -cx | llm -m claude-fable-5.1 -s 'animate this' Here's the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... It's excellent!

Now that it's a solved benchmark, can we get the 3d animated version?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#733
post #712

Earlier quoted context omitted.

From Artificial Analysis cost per task, it looks like Fable 5.1 (max) is more expensive per task than Fable 5 (max)? Cache hit price went down, but the other components still add up to more. Edit: 5.1-xhigh seems to be cheaper than 5-max, and 5.1-xhigh has a higher index score than 5-max. Also interesting that Fable 5.1 (high) is comparable to Opus 5 (max), but nearly half the price. https://artificialanalysis.ai/mod…

From my limited testing of just 2 hours, reasoning output of 5.1-max is at least 7x of 5-max, on the same project and comparable prompts. It reasoned for ~2 minutes trying to figure out an appropriate directory name. I've never seen 5-max do that. Could be a misconfiguration though.

I haven't dug deep but I burned through 30% of my weekly usage in a few hours which shocked me at first.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#735

Earlier quoted context omitted.

I sometimes have that feeling too, then ask another LLM to do a code and vulnerability review and OMG: rookie mistakes, over complications and security gaps even a 1st year student would not make regularly. So.. one more year of untreated bipolar AI psychosis I guess..

I think these kinds of comments really need to say which LLM that is. There's an enormous difference in skill between the frontier ones and say the Google search AI.

Codex Luna, Terra and Sol. Claude Opus, Sonnet and sometime Fable.

They all work, they all are "good", they all are both "smart" and commit incredible basic mistakes a fair amount of times.

Then there's the cost situation..

Re: Claude Fable 5.1 and Claude Mythos 5.1

#738

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

> * Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.

That wasn’t Anthropic. Clearly not a well informed take.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#739

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Does the new writing style now have EU level watermarks?

It does, as per 'Compliance with the EU AI Act' section.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#740
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

Maybe you could show a side-by-side comparison of pelican images. One image doesn't really make the improvement clear for someone like me. That would be a great help.
Post reply on HN