Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

601–610 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#601

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

Does it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"? "Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure. "Fail closed" is the opposite -- system has power and is live. Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and compu…

Nah, "fail open/closed" means that in failure mode something is open. It's "good" when something is a circuit and what failed is a fuse, but it's "bad" when it's your API security. If it's a valve, it probably can be good or bad depending on the use case.

It doesn't mean "fail open" is always the desired/safe outcome. It goes back to 1872 air brakes on a train. The goal is to "fail in safe mode", sometimes it's open, sometimes it's closed.

From the top of my head, where "fail open" is the desired outcome:

- emergency doors

- industrial cooling

- pressure valves

- probably something in HVAC

Note that none of these are "computer security people".

Re: Claude Fable 5.1 and Claude Mythos 5.1

#602

Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke. What they have done: * Nerfed Fable, as many of noted it's useless * Leverage Mythos as a marketing strategy, claiming its too good to release * Removed thought traces, one of the only useful things to make sure your prompts are working correctly…

> Nerfed Fable, as many of noted it's useless I certainly don't take AI advice from HN, but this is amazing . Useless? Yes, the safeguards are ridiculous and obnoxious, though I can say that 5.1 greatly relaxes them (just doing a hardening of a project parallel with this comment, which 5.0 refused to do...so did Sol and Gemini, fwiw. The Gemini one is a laugh, because 3.1 pretending like it's a dangerous tool is simp…

[flagged]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#603
post #405

I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets. I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited. I urge Anthropic to get be…

Have you tried Grok 4.6, if you're focused on token budgets? In a league of it's own for tokens/intelligence.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#604
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

A pelican is a bird, not a person. The knees bend the other way.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#605
post #365

What I don't see in the comments: "I had a specific problem I couldn't solve with the previous version of this LLM. But the improvements in this version unlocked the solution for me." What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too. I use coding agents. To me they are very useful. But what I spend on them isn't…

We rarely upgrade our phones or MacBooks because the newer version can do something the previous one literally couldn’t. Often it’s the efficiency, speed, battery life, etc, combined, that lets us push the hardware further. I get your point, but we can only have groundbreaking leaps once in a blue moon. That doesn’t mean incremental improvements aren’t useful.

What you were describing our products at the top or near the top of their S curve. That only works if a product has achieved a mature market that's big enough to sustain further product development. Apple might take a percentage point of market share from Windows, and Linux might take a 10th of a point, but nobody is suddenly going to find, or lose, a big chunk of the market.

The problem frontier LLMs face is that they are hundreds of billions to trillions of dollars short of finding that market that's big enough to sustain capex commitments and further product development. If they don't find something groundbreaking, they are going to have a very painful year next year, maybe even starting this year for some of them and their data center partners.

Anthropic and OpenAI can't afford to live in a world where LLMs are at or near the top of their S curve.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#607

Earlier quoted context omitted.

> Nerfed Fable, as many of noted it's useless I certainly don't take AI advice from HN, but this is amazing . Useless? Yes, the safeguards are ridiculous and obnoxious, though I can say that 5.1 greatly relaxes them (just doing a hardening of a project parallel with this comment, which 5.0 refused to do...so did Sol and Gemini, fwiw. The Gemini one is a laugh, because 3.1 pretending like it's a dangerous tool is simp…

[flagged]

That's, uh, a great contribution. Thanks. It's super important that HN learns how this sounds like to you.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#608

Earlier quoted context omitted.

I'm legitimately out of the loop; what is going on/broken with Opus 5?

it just doesn't interact good with human beings, and it leaves incredibly strange long winded comments within code filled with session context that will likely not be relevant later on. Also always seems to have this annoying tendency to leave "questions for you" at the bottom of every output. Just a high friction human interaction type model, imo should never have even been released, regardless if it scores better o…

I have to wonder if everyone else is just running these models raw without any custom instructions. I hear all these things about voice and code comments and those are all things I've dealt with long ago via claude.md instructions, rules, and hooks. My claude can already respond in any "voice" I want and the quantity and quality of comments is within my control.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#609
post #191
post #35

“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.” I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better. I don’t think Anthropic realizes that…

I just can't stand how often Claude says something like "And the honest part? It's..." Like, were the other parts not honest? I don't understand how Anthropic let it get like this, it's been such a clear regression

[dead]

Re: Claude Fable 5.1 and Claude Mythos 5.1

#610
post #452

All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.

I'm a kernel engineer. Fable 5 refused all my requests, falling back to Opus 4.8. My wife is a chemist. Her experience wasn't much better.

I'm curious about this because I've had Fable decompile games and help me understand what's going on inside the game itself and it never complained. I'm not sure what it takes to trip the "safety" guards but digging into game code and data files doesn't seem to be a barrier at all. I've used CC to build some personal game mods a few times now. Once for a game with no modding capability explicitly exposed.
Post reply on HN