Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

531–540 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#531

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting t…

That’s actually common. Not in academia but a lot of enterprises are specifically not using Fable because Anthropic doesn’t provide a Zero Data Retention mode like they do for Opus. Even at my employer when Fable is available, some employees just aren’t comfortable using it when they perceive that they are working on extremely sensitive research.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#532
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser... You're no longer allowed to edit the conte…

Interesting, I liked to experiment with a second model "simplifying" and summarizing the previous messages and continue.

Needless to say, it improved output on following messages by whatever metric I cared for.

Not sure why would they prevent it.

I give you a chain of messages, what do you care for what the origin is?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#533
I'm afraid watermarking could restrict applications where LLMs can be safely used to assist with writing. If I write something myself and use an LLM to proofread it, without watermarking I can confidently say that corrections done by LLMs are small and insignificant enough to claim that the text is still authored by me, not by the model. With watermarking, however, I will never be sure if the result will not be flagged as AI generated, even if the AI contribution is very minor.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#534
post #19
post #2

This time it came with a usage reset

Great, my usage reset is in 10 hours ... And my 5 hour window was due to be reset in 2 hours (barely used), now its in 5 hours - so this reset effectively gives me 1 less 5 hour reset for this weekly cycle.

Unless you run overnight, you could schedule a cron job to send a basic claude -p prompt such as "reply with hello" using haiku to align your usage windows. That's what I do.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#535
post #479

My main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches" .. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd […

> Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers"... It's copywriting. They fed these models the internet, which is loaded with it.

And turns out, the "frontier" labs have no human oversight of the training data going into these models... Explains so much

Re: Claude Fable 5.1 and Claude Mythos 5.1

#536
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

To be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.

I think this whole distillation argument is between fully overblown and bogus.

In any case, highly misunderstood.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#537
post #479

My main gripe with LLMs is the cringe AI phrasings that they use in UI elements. Pompous things like "Your keys, supercharged" or weird yoda-speak stuff like "searches the app remembers" instead of just naming the thing "Learned searches" .. you know, proper GUI copy like it was done for the past decades. I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd […

> Where are all these verbal tics coming from and why is it so hard to get rid of them?

It’s a side effect of post-training for effectiveness and efficiency at technical tasks.

Over time the models learn to pack as much information as possible into their available context window, because that’s one way to increase the effective intelligence.

Humans do this too with industry jargon, dense tech-talk, etc.

We have a limited capacity so packing it densely maximises what we can do with it.

If you’ve ever heard a “non technical” manager complain about the terminology in an IT meeting — this is why.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#538
The safeguards and required extra retention is still not gone. Further more they are working to create separate tiers of access with the new biology program instead of giving everyone equal access to AI. Anthropic once again are showing they can not be trusted.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#540
post #502

According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens. This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency. Fable 5: https://artificialanalysis.ai…

https://artificialanalysis.ai/models/claude-fable-5-1-high

On high it gets the same score as 5 with max effort while costing only half as much.

Post reply on HN