Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…
[flagged]
Claude Fable 5.1 and Claude Mythos 5.1
371–380 of 1001 posts
Re: Claude Fable 5.1 and Claude Mythos 5.1
#372Fable 5.1 is actually more expensive than 5.0 when run on the Artificial Analysis suite: https://artificialanalysis.ai/#intelligence-efficiency-tabs
Re: Claude Fable 5.1 and Claude Mythos 5.1
#373(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
Do you know if Opus 5.1 is coming and will have improvements in writing style too?
Re: Claude Fable 5.1 and Claude Mythos 5.1
#374Earlier quoted context omitted.
Does it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"? "Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure. "Fail closed" is the opposite -- system has power and is live. Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and compu…
That doesn't make sense at all. Fail open means the method of it's use is still in use. Say you have a door that has powered locks. You want it to fail "open" so that when the power goes out, it's still useable, and people can get out. That's the source of the term.
The concept goes back to a pressure cooker invented in 1679 by Papin.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#375(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#376(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…
I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…
Give me TERSE.
Re: Claude Fable 5.1 and Claude Mythos 5.1
#377"Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards." Can't believe they haven't at least figured out better messaging. If we take them at their word, it's hard not to read it as a messiah complex, that they think they're the only ones capable or worthy of making these decisions. I don't believe them, but I wouldn't be surprised if the articulated…
Can't say I had such troubles actually, no. Their position can be extended to any and every model provider just fine, it does not single them out specifically.
Surely there's a less hyperbolic and ad hominem-y way to take issue with this? I don't think following up a critique about ineffective messaging with one centered around a demagogue reach is particularly compelling at least.
Their argument is that the model provider owns the safety story, and that as such, they consider the extraction of capabilities (which washes the guardrails) as a failure on their side. If this makes you think of personality traits, I'm not sure you're engaging with their position earnestly. It most certainly doesn't leave me any more equipped to disagree with them either.
If you instead highlighted how awfully convenient it is, however...
Re: Claude Fable 5.1 and Claude Mythos 5.1
#378The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#379What I don't see in the comments: "I had a specific problem I couldn't solve with the previous version of this LLM. But the improvements in this version unlocked the solution for me." What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too. I use coding agents. To me they are very useful. But what I spend on them isn't…
Re: Claude Fable 5.1 and Claude Mythos 5.1
#380> We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. I'm not an emdash hater but this isn't how you use them. It should be a comma.
Em dashes are commonly used to add emphasis, even where you would ordinarily use a comma. Their flexibility is why many people love them! See https://www.merriam-webster.com/grammar/em-dash-en-dash-how-...