(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting t…
Claude Fable 5.1 and Claude Mythos 5.1
531–540 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#532Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…
These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser... You're no longer allowed to edit the conte…
Needless to say, it improved output on following messages by whatever metric I cared for.
Not sure why would they prevent it.
I give you a chain of messages, what do you care for what the origin is?
Re: Claude Fable 5.1 and Claude Mythos 5.1
#533Re: Claude Fable 5.1 and Claude Mythos 5.1
#534This time it came with a usage reset
Great, my usage reset is in 10 hours ... And my 5 hour window was due to be reset in 2 hours (barely used), now its in 5 hours - so this reset effectively gives me 1 less 5 hour reset for this weekly cycle.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#535My main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches" .. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd […
> Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers"... It's copywriting. They fed these models the internet, which is loaded with it.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#536Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…
To be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.
In any case, highly misunderstood.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#537My main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches" .. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd […
It’s a side effect of post-training for effectiveness and efficiency at technical tasks.
Over time the models learn to pack as much information as possible into their available context window, because that’s one way to increase the effective intelligence.
Humans do this too with industry jargon, dense tech-talk, etc.
We have a limited capacity so packing it densely maximises what we can do with it.
If you’ve ever heard a “non technical” manager complain about the terminology in an IT meeting — this is why.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#538Re: Claude Fable 5.1 and Claude Mythos 5.1
#539Re: Claude Fable 5.1 and Claude Mythos 5.1
#540According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens. This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency. Fable 5: https://artificialanalysis.ai…
On high it gets the same score as 5 with max effort while costing only half as much.