Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…
Claude Fable 5.1 and Claude Mythos 5.1
801–810 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#802Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…
That's not part of the EU regulations. You only need to say that it is created by AI, and then only under certain conditions.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#803Earlier quoted context omitted.
We will have more bugs. Even the best models with the best software engineers will produce bugs. There are two reasons : first the pressure to produce more and second LLMs will always produce slop
>> LLMs will always produce slop Such a low-quality comment
Re: Claude Fable 5.1 and Claude Mythos 5.1
#804Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…
* Letting you Sign-Up-with-Apple on iOS but not Sign-In-with-Apple on web, but supporting Sign-In-with-Google
* Not letting you remove your payment info
* Not letting you change your email
* Seemingly no way to get real support
Re: Claude Fable 5.1 and Claude Mythos 5.1
#805I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…
My personal guess is that it's one of those. With effective context engineering it's hard to use all 20x quota, the limit becomes your own attention and time really.
You may argue that you're doing multiple, parallel extreme effort tasks – which may be true but then again, there will be results to actually look at sooner or later and that takes time.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#806The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…
It's because GPT Sol is equally good and established a price ceiling
Re: Claude Fable 5.1 and Claude Mythos 5.1
#807Re: Claude Fable 5.1 and Claude Mythos 5.1
#808The bar has dropped when the dialogue evolves to "output is less annoying" rather than some interesting new capability
That doesn't make them less impressive, it just means people are shifting their focus more towards their own day to day experience with these things because we're relying on them so much now.
Like when the novelty of the automobile wore off, I'm sure people were starting to say "it's a bumpy ride though, isn't it?"
Re: Claude Fable 5.1 and Claude Mythos 5.1
#809(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
Not sure if something like this is already on the table, but I would like to see Claude responses more in line with Simplified Technical English [1] by default. I find those writing styles a lot easier to read. This has been standardised as ASD-STE100 [2]. I've seen few people making SKILL.md files with that in mind, which works great, but having this by default without invoking the skill command would be better. 1.…
As a workhorse, GLM is so good, but goodness, its prose, wherever needed, makes me feel like going to a park and kicking all the benches there endlessly. And it doesn't change!