Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

551–560 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#551

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

100% convinced their raw output is intended as further inputs, and my workflows have been comfortable and efficient treating it as such. If you really need to read slop, you ask your agent to give it to you in a style that works for you. I can imagine a world where the slop from others doesn’t hit us directly but gets personal mediation.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#552
post #510

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

What do you mean fable is useless?

(not op) It cannot be used to develop applications. Every application needs to be secure in some way, and any such mention in a review triggers Fable's upsell feature.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#553

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

Text watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b... .

This is hilarious this keeps being repeated by the true believers ad nauseam.

Also, don't apply EU law to the world. It's a knee jerk reactionary regulation by a bunch of aging ding dongs that can't print their emails.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#554

Earlier quoted context omitted.

I hope that this also applies to the Subscription usage. As that can then stretch out Fable usage by a lot more.

My understanding is subscription usage generally has free cache reads, but I'm not sure if maybe Fable was different in that regard.

The fact that we don't know is part of the problem. Subscription usage has always been pretty opaque.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#555

> This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their orga…

It does change the output, they never said it did not. They said it would not _noticeably_ affect performance.

You have a misunderstanding. Watermarking does not bias the responses in any way. How is this possible?

Before: "He leaped at the chance" - 33%. "Jumped at the opportunity" - 66%.

After: "He leaped at the chance" - 33%. "Jumped at the opportunity" - 66%.

But if you refresh your response from Anthropic 100 times:

Before: "Jumped at the opportunity" He leaped at the chance" "Jumped at the opportunity"

After: "He leaped at the chance" "He leaped at the chance" "He leaped at the chance"

The second one is detectable as being watermarked.

davmre has a good explanation that's more in-depth.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#556

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Fable is useless. Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company." Fable: "yeah, no."

It makes sense. Even if it finds some exploit on your own code, who's to say you can't reuse the same exploit on some other system?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#558

Earlier quoted context omitted.

As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.

> As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The same way it did in the previous versions: brute force. I don't believe that LLMs have any particular intelligence we don't, but there's an endless list of problems we either don't have bodies to throw at, or the bodies we can throw at it, don't have such…

> They will not revolutionize human knowledge, but they can definitely widen it a lot.

I am generally quite enthusiastic about all this, but my biggest fear is that we will not recognize the extreme need for more scientists at a time when there is so much more science to be done. The rate of scientific understanding must keep pace with the amount of science being output, both for verification and further discovery. It's a pipelining issue, and I predict a stall in the bits that require the (currently rare) people who know what they're doing.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#559

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

My hunch is that much of the model tuning to make it more effective has been for its internal thinking prose. That leaks out into its external writing prose.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#560

I'm confused about Anthropic's pricing. Can anyone explain why Sonner 5 is $2/MTok in and Sonnet 4.6 is still $3?

They originally released it at a "temporary discounted price", then made it permanent (probably due to competitive pressure). It's still way more expensive per task, due to tokenizer changes and general verbosity.
Post reply on HN