Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

451–460 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#451

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

  > They're packing lots of signal into fewer words 
FYI, these are so-called `load-bearing` words.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#452

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#453
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

"Humans have a token limit too" - that's so good and it explains so much of the fatigue that myself and colleagues/peers have about Claude in particular.

I think it's not just token limits - I think it's because it's so _dense_.

You get a week of research and debugging and testing compressed into a few pages. Even if it's explained well, it's just so much information. And since it's AI, I'm constantly second guessing "is that really true?" and it's exhausting.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#454
post #390

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

The only company to use Claude.md instead of Agents.md standard

With new watermarking you may now get Hullaballooing.md

Re: Claude Fable 5.1 and Claude Mythos 5.1

#456
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

Do you think model trainers are pelicanmaxxing now?

There was a real probe into this, seems like no:

https://dylancastillo.co/posts/pelicanmaxxing.html

https://news.ycombinator.com/item?id=49010129

Re: Claude Fable 5.1 and Claude Mythos 5.1

#457

Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would: 1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan 2) either fix opus 5, make it completely free, or delete it entirely

I'm legitimately out of the loop; what is going on/broken with Opus 5?

try sol and you'll see

Re: Claude Fable 5.1 and Claude Mythos 5.1

#458

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

Cut them a break. They are trying to IPO soon.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#459

> We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. I'm not an emdash hater but this isn't how you use them. It should be a comma.

Grammatically an emdash is fine in most places a comma is fine. It adds a bit more emphasis to the bit after the dash. I went to the grocery store, and bought tomatoes. I went to the grocery store---and bought a Ferrari. The second one has a bit more of a dramatic pause. "Eats, Shoots, and Leaves" is a fun book with a great chapter about the dash with many good examples.

Emdashes and commas aren't interchangeable, and your example there demonstrates one great reason why. The emdash establishes a discontinuity rather than one thing flowing into another, which is why the tomatoes don't merit one but the Ferrari does: you are using the emdash to emphasize the situational irony.

Going back to Anthropic's post:

> They’re the world’s most advanced models for coding and knowledge work---and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.

The first thing directly implies and flows smoothly into the next---or would, if not for the awkward emdash. There is no discontinuity, no twist or shift in context, no implied question and provided answer, no punchline. It's just distracting.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#460
"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%."

They show this off, but artificial analysis contradicts the statement. Fable 5 cost $3.14 per task, while 5.1 cost $3.69 -- around a 15% jump in pricing.

https://artificialanalysis.ai/

These, IMO, are marginal improvements for a more expensive model. I stopped using Claude ~3 months back; its outputs are too jargoned, it makes architectural decisions that are not right, and it's incredibly pricey for what it is. Each decision it makes, it acts as if a problem as major as world hunger has been solved. And the overly verbose code comments, strange commit descriptions, duplicate code, and slop it generates -- which I know is not specific to Fable -- is just too much for me.

I found the best is to use something like Deepseek V4 Flash -- with a fast TPS provider -- and work on the code myself. For agentic work with computer use, GLM 5.3 flash with Hermes Desktop works well.

Post reply on HN