Earlier quoted context omitted.
I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.
just buy two 20x, no?
Claude Fable 5.1 and Claude Mythos 5.1
771–780 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#772Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…
Maybe you could show a side-by-side comparison of pelican images. One image doesn't really make the improvement clear for someone like me. That would be a great help.
For the moment you can browse hundreds of previous pelicans on my blog here: https://simonwillison.net/tags/pelican-riding-a-bicycle/
Re: Claude Fable 5.1 and Claude Mythos 5.1
#773(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
1. https://en.wikipedia.org/wiki/Simplified_Technical_English
Re: Claude Fable 5.1 and Claude Mythos 5.1
#774Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#775(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#776What I don't see in the comments: "I had a specific problem I couldn't solve with the previous version of this LLM. But the improvements in this version unlocked the solution for me." What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too. I use coding agents. To me they are very useful. But what I spend on them isn't…
I had two sessions this morning that prior fable and sol sessions were stuck on, where iterations just resulted in _different_ bugs. (One kind of tricky fe layout problem, the other was a backend refactoring that was complicated by trying to aggregate a couple prior sessions that crashed). I summarized each into new fable 5.1 sessions, and both seem to have arrived at reasonable solutions that only need a few nits re…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#777> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#778My main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches" .. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd […
This style provides a high entropy basis distribution, so they can from a bigger pool to pick from and phrases to watermark the sentence.
You have much bigger variation of this idiotic phrases and words, which states a simple fact in that sophisticated and twisted manner.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#779To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. T…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#780The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…
I'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything. Optimizing a OS build? -> block Securing a container -> block 60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked h…