Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

241–250 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#241
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

In a way, this is the only benchmark I care about now. :)

Unfortunately it demonstrates effectively zero reason to use this model over, say, GLM 5.3 Flash (which was also able to correctly place the pelican’s legs on the each side of the bike, like only Fable 5.1 xhigh was able to do here)

I still enjoy seeing the pelicans.

Edit: Ok, max effort made a darn good pelican.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#242
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

If you ask any model to write as tables to enumerate points, and BDD for logical flows, it’s like 50x less strain on you

Re: Claude Fable 5.1 and Claude Mythos 5.1

#243
I've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem.

In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm busy on other tasks. It also has extensive access to my computer, other computers on my network, my internet. It's really helpful when you give it a lot of resources, but right now I have very autonomous, very smart agent running around more or less unattended with a lot of resources.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#244

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Qwen is all you need.

Not coincidentally, Claude is all Qwen needs.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#245
post #143

"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations" Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??

Makes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#246
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

I've developed a habit of adding into my prompts "please keep your response concise and succinct" or "I'm trying to cram, please only provide the minimum level of technical detail necessary to understand this topic" I find it helps immensely but it'd be nice if I didn't have to do that.

why so many people add 'please' when asking machine to do something? Was there actually research that when you SCREAM or curse it follows your instructions better?

P.S. Although my wife insists that I should stay polite in case AI overlords remember how I treat them ...

Re: Claude Fable 5.1 and Claude Mythos 5.1

#247

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

Less frequent context truncation, too, leading to better scores?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#248
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

Now that it's a solved benchmark, can we get the animated version?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#249
post #191
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

I just can't stand how often Claude says something like "And the honest part? It's..." Like, were the other parts not honest? I don't understand how Anthropic let it get like this, it's been such a clear regression

i think they took a huge bet that speaking like a ted talk was going to be a vast popular differentiator in their offering, i don't think they anticipated that people were going to make fun of it, that it could become a meme..that it could get in the way of getting stuff done and result in cancellations.

it's downright exhausting to read claude, the language style was a regression imo.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#250

Earlier quoted context omitted.

Fable is useless. Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company." Fable: "yeah, no."

It isn't exactly hard for a bad actor to come up with that prompt

well no crap right? Except I submitted for an exception, even sending my linkedin and using a company email address. it should be extraordinarily obvious we own this code.
Post reply on HN