Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

201–210 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#201

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

> It sounds a lot less stereotypically like other Claude models Don't give me hope. I've strained eye muscles from rolling my eyes so hard every day at how Claude writes. Edit: first discussion with Fable 5.1 "This is the right question and it needs a real trace, not a guess." Sigh.

[dead]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#202

Earlier quoted context omitted.

Please bring to the other models, and also please only apply the AI text watermarking only to EU citizens. I may not be able to tell when Claude writes about things i don't know, but in CC it writes about my code and it is obvious.

You’re probably better off organizing a campaign to pressure Congress to prohibit American corporations imposing foreign laws on Americans, which is what this text watermarking is, regardless of how you feel about it. I think it’s a precedent we really don’t want to go down if you believe in democracy and self-determination. It also clearly establishes or the very least moves in the direction that you don’t actually…

I think Congresspeople hearing that EU AI Act is forcing secret codes into the infrastructure of American technology across all industries is sufficient.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#203

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

I notably had an issue that it wouldn't work on a "remote execution" (running a command over SSH) coding problem until I did a sed to remove the word "execution". Incredibly dumb. I'm not doing any murders. Easiest to just switch to the Chinese models.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#204

"Cache reads now cost 75% less, or $0.25 per million tokens." For me, at a typical 95% cache hit rate, I think my optimal context window size before autocompaction goes from ~200K to ~400K tokens. Great for longer horizon tasks.

looks like it is only for api.....

may i ask where did you get this?

i try to look through the docs, but i didn't find where they said its only for API

is it in the system card?

really hope not, that change the only positive part in this release

Re: Claude Fable 5.1 and Claude Mythos 5.1

#205

Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely

I'm legitimately out of the loop; what is going on/broken with Opus 5?

it just doesn't interact good with human beings, and it leaves incredibly strange long winded comments within code filled with session context that will likely not be relevant later on.

Also always seems to have this annoying tendency to leave "questions for you" at the bottom of every output.

Just a high friction human interaction type model, imo should never have even been released, regardless if it scores better on whatever tests, its a horrible experience and a downgrade over past models.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#206

Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely

I'm legitimately out of the loop; what is going on/broken with Opus 5?

They've nerfed a bunch of models, especially Opus 5. Nobody knows why, but overall things have gone downhill significantly.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#207

Earlier quoted context omitted.

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> They're packing lots of signal into fewer words There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are…

It has always seemed to me that they're hacking for dopamine response in moderately interested data labelers.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#208
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

One thing I've noticed and HATE, is that when you increase thinking-effort, that seemingly increases response-length. Meaning that X.High is longer than High, which is longer than Medium, etc. Which is kind of the inverse of how people work; a really smart person can condense difficult ideas into simple[r] terms. Whereas people who struggle speak a lot but say very little. High/X.High do seem to deliver better qualit…

The higher the effort the more things Claude checks, and it's eager to tell you about all of them

See, this insight it had early on looked like a red hering for a while, but then turned out to be load-bearing. And that's not just a difference in semantics, it changed the whole conclusion (spoiler: it didn't). And Claude is very eager to tell you about this exciting journey

Re: Claude Fable 5.1 and Claude Mythos 5.1

#209
post #46

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from o…

To be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.

I have a hard time believing whatever prompts get Claude to reason can stay relevant secret sauce for long anyways. It’s not hard to A/B test something that gets you close enough, and it’s not Ike anthropic has uncovered the global optima of reasoning prompts.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#210
post #4

Bit of a discount if you're using caching: > same input and output prices, with cache reads at a quarter of the cost This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.

And yet, despite this, the quota limits went down by 17%.

In my opinion, this is a bit disingenuous.

They were _temporarily_ increased in May by 50% [1]. They continued to extend them through July and August (admittedly, their messaging around this has just been a complete mess and they frequently pushed the deadline back as it approached).

So, now they are giving you a 25% quota increase compared to where things originally stood in May.

So, let me ask you this: assuming you knew that the 50% quota increase was temporary all along, would you then have complained about Anthropic restoring things back to the original limit?

[1] https://www.anthropic.com/news/higher-limits-spacex

Post reply on HN