Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

761–770 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#762

To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. T…

I agree that intelligence at cost is exciting right now, especially if you view AI as a tool. The clock is ticking on subsidized tokens and cheap or free local inference will be the future. I think people want AI to be an oracle for prediction and discovery, which is where the sota models come in. But each release seems more iterative and underwhelming than the last. When the latest models regularly reveal unexpected insights, like how to get my execs to stop demanding hand-wavey 10x productivity gains, somebody let me know.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#763

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> the models writing more for themselves and each other than for humans

What does this means?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#765
post #747

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

You have a serious engineering problem if you're not able to find the source of a crash after years.

If it’s rare and the impact is low, then it’s not getting prioritized. It doesn’t matter how much time passes if you decide not to spend time investigating.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#766
post #747

Earlier quoted context omitted.

You have a serious engineering problem if you're not able to find the source of a crash after years.

You are either seriously naive, or have never worked on any large and complex legacy codebase.

I’ve been living in a bubble with my .NET day-job, where debugging/tracing/postmortems are a breeze. Compare with, say, a CORBA or DCOM system, deployed to prod with uber-optimized binaries without any debugging-symbols.

So it’s not that I haven’t worked on large-scale, complex legacy systems - but that I haven’t worked on any large-scale, complex legacy systems written in languages bereft of runtime reflection and verbose error reporting.

—————

It’s also possible that the bug was never found because its impact was so minimal: e.g. 1 crash per year, each causing 3 minutes’ downtime in a noncritical system: that’s something that will never get investigated fully.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#767

On both my work (Team Premium) and personal accounts (Max 20x), Fable 5.1 hit the 5-hour limit before it could finish the first task I gave it. On my work account, it took about 30 minutes, and on my personal account, less than an hour. This has never happened to me before, but if this is normal behavior, Fable 5.1 is essentially unusable.

How did you hit the 5-hour limit if it took one hour?

Re: Claude Fable 5.1 and Claude Mythos 5.1

#768

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

I sometimes have that feeling too, then ask another LLM to do a code and vulnerability review and OMG: rookie mistakes, over complications and security gaps even a 1st year student would not make regularly. So.. one more year of untreated bipolar AI psychosis I guess..

At least we are at a point where we can have AI review code and reliably find real problems. That alone is incredibly valuable.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#770

Earlier quoted context omitted.

I literally only make it halfway through the week until my weekly usage runs out. This is using only Opus, no fable, and I'm on the max x20 plan. It's become ridiculous.

just buy two 20x, no?

or switch to codex
Post reply on HN