Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

931–940 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#932

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

[flagged]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#933
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

[flagged]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#934

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> Opus prose style/smell we all have grown weary of

I bet everyone will grow wear of absolutely any style a stochastic parrot would use continuously ad nauseam. The lack of human variability is the reason, not the style itself.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#935
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser... You're no longer allowed to edit the conte…

How do they still purport to champion alignment and explainablity of AI if the reasoning traces are going to be hidden.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#937
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

Anthropic:

AI models should be explainable so that we can ensure and verify alignment. Responsible AI 101.

Also Anthropic:

No not like that.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#938

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.

For one, by synthesizing the results of multiple papers and suggesting novel experiments. If one paper sets constraints X for some system, and another paper sets constraints Y where Y!=X for a system that's similar but slightly different, then that's fertile ground for an experiment that can extract the more general underlying principles. This has already happened for domain-specific AI in fact, but the idea here is that it will become routine with general AI systems, as is happening now with math.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#939
post #304
post #248

Earlier quoted context omitted.

Now that it's a solved benchmark, can we get the animated version?

I didn't want to shell out for Max again, so I piped the SVG created by Max back into Fable 5.1 at its default thinking level (of high): llm logs -cx | llm -m claude-fable-5.1 -s 'animate this' Here's the result, which cost $1.37: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... It's excellent!

Dang the little fish jiggling around the basket is a nice touch

Re: Claude Fable 5.1 and Claude Mythos 5.1

#940

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I recently ended my Claude subscription, returning back to ChatGPT because I could no longer bear Claude’s prose, finding it excessively verbose, robotic, repetitive, and condescending.
Post reply on HN