Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

621–630 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#621
post #477

Earlier quoted context omitted.

Text watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b... .

> Text watermarking has no effect on output quality It has an effect, and it's negative. It's hoped that the effect is negligible, and it probably is, but the whole point is that it has an effect.

I am pretty sure they did A/B testing to show it didn't. I could gave sworn they even released a quiz were the user has to try and guess which answer is watermarked or not and it was impossible to tell.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#622
post #477

Earlier quoted context omitted.

Text watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b... .

> Text watermarking has no effect on output quality It has an effect, and it's negative. It's hoped that the effect is negligible, and it probably is, but the whole point is that it has an effect.

Its essentially swapping out the psuedo random number generated with a differently seeded one iirc.

It has an effect on the output, but not the output quality

Re: Claude Fable 5.1 and Claude Mythos 5.1

#623

Earlier quoted context omitted.

As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.

> As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The same way it did in the previous versions: brute force. I don't believe that LLMs have any particular intelligence we don't, but there's an endless list of problems we either don't have bodies to throw at, or the bodies we can throw at it, don't have such…

We are not limited by intelligence, or bodies.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#625
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

Yes! I have my .md's have

"If you respond with more than 3 paragraphs, give me a TLDR"

"Do not assume I know all technical jargon, please explain things plainly"

Re: Claude Fable 5.1 and Claude Mythos 5.1

#626
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

Just the other way I was thinking that if I asked "What does Lamborghini do?" the only correct way to answer is a single sentence "Which Lamborghini are you referring to?". But LLMs will fail at this question: they will tell you about Lamborghini's latest car and mix some history in it. Just try. Which is the wrong answer anyway, because there's at least two major companies called Lamborghini, one making cars, one ma…

This is ... unnecessarily pedantic. Anybody in my social universe who asked me that question would undoubtedly expect "they make cars".

If you're picking nits, why not focus on the word "do" and (wrongly) expect an answer like "Lamborghini (either of the two main companies of that name) does not 'do' anything - the companies employ humans who 'do' things. Lamborghini is a legal entity established to allow humans to 'do' things, such as make cars, or agricultural equipment."

Shared context is a thing. Reducing every conversation to first principles is not always required. Get a grip.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#627

The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general. Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is h…

I'm a heavy user and fable is great the #1 reason I stopped using it was the horrible safegaurd filter. I found sol close enough in capability and have only been blocked when my request was an obvious offensive cyber work. Fable blocked me on almost everything. Optimizing a OS build? -> block Securing a container -> block 60% is nowhere near enough for that safegaurd system. This just means I am going to be blocked h…

It dramatically improved about a week ago - most of my blocked projects are now completely functional.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#628

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

We will have more bugs. Even the best models with the best software engineers will produce bugs. There are two reasons : first the pressure to produce more and second LLMs will always produce slop

>> LLMs will always produce slop

Such a low-quality comment

Re: Claude Fable 5.1 and Claude Mythos 5.1

#630

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

claudism really sucks, Gemini and codex output so much better, way more like a real human being.
Post reply on HN