This has never happened to me before, but if this is normal behavior, Fable 5.1 is essentially unusable.
Claude Fable 5.1 and Claude Mythos 5.1
311–320 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#312Why aren’t these models available on subscription plans? I tried the old fable and it didn’t seem worth paying for. It still made errors like Opus does so I might as well use the included model…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#313Earlier quoted context omitted.
One thing I've noticed and HATE, is that when you increase thinking-effort, that seemingly increases response-length. Meaning that X.High is longer than High, which is longer than Medium, etc. Which is kind of the inverse of how people work; a really smart person can condense difficult ideas into simple[r] terms. Whereas people who struggle speak a lot but say very little. High/X.High do seem to deliver better qualit…
With LLMs, you're still mostly read things "off the tip of the tongue". A better comparison is observing a smart person talking to themselves while working on a tough problem. EDIT: also there's a reason the dial is called "effort", not "smarts".
Re: Claude Fable 5.1 and Claude Mythos 5.1
#314Re: Claude Fable 5.1 and Claude Mythos 5.1
#315(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#316Earlier quoted context omitted.
I've developed a habit of adding into my prompts "please keep your response concise and succinct" or "I'm trying to cram, please only provide the minimum level of technical detail necessary to understand this topic" I find it helps immensely but it'd be nice if I didn't have to do that.
why so many people add 'please' when asking machine to do something? Was there actually research that when you SCREAM or curse it follows your instructions better? P.S. Although my wife insists that I should stay polite in case AI overlords remember how I treat them ...
Re: Claude Fable 5.1 and Claude Mythos 5.1
#317I am absolutely thrilled that they reset weekly limits. I have been experimenting with highly autonomous work (5+ hours continuous) and fable seems excellent at this, especially when using subagents. I ran out of Fable capacity and was bummed out that my experiment would take longer to complete. Now I'm super happy I get to continue it
Re: Claude Fable 5.1 and Claude Mythos 5.1
#318Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#319The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…
It wouldn't surprise me if we start to see minimal performance gains from incremental changes to base models. It seems like the gains from the Opus 4.5+ incremental updates were a result of Anthropic learning a lot about post-training, the gains from RLVR, etc.
If new post-training techniques are seeing diminishing returns, we could just be back to waiting for new large pretraining runs at larger sizes for gains (even if those ultimately end up getting distilled down into smaller models because the economics for serving anything larger than Fable isn't practical).