Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

741–750 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#741
post #144

Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht... I'm still waiting for effort max to finish. EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Took just under…

Maybe you could show a side-by-side comparison of pelican images. One image doesn't really make the improvement clear for someone like me. That would be a great help.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#742
I have to say, I am quite frustrated with Anthropic lately. I so badly want to use Fable to work on a side project of mine, which I used to do previously with no issues. But lately, they must have made some classifier change because it keeps hitting their stupid, overly-hyper-aggressive safeguard due to 'general_harms'.

Guys, listen to your feedback please. I hadn't used OpenAI products in quite a while until this issue came around. They seem to have MUCH smarter safeguards than Anthropic does.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#743
post #731

Fable is too expensive for general use I think this is why it hasn’t received as much attention as expected since Fable came out Developers always work while trying to find ways to work continuously for a 5-hour session without disconnecting. Fable has had the experience of using up all its tokens before I even realized it because the burn rate was too fast. Since then, I always use only Opus. For Fable to become a c…

Fable is much more expensive both in time and tokens for a marginal increase in productivity.

I find Fable can solve in minute things that Opus struggles with. Of course Fable can struggle too.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#744

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

this sounds very much correct and i don't really mind it for that reason. i do a lot of long-running tasks and i feel like it can really pick up on its own thread easier if i just let it write in its own way.

i am also using Opus for a hobby teaching agent, and the way it writes the prompts is "cringy" but they seem to work well. i almost want it to continue doing this internally, it understands best this way.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#745
post #467

Earlier quoted context omitted.

While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting t…

Isn't it showing a problem with an academia? "I don't want to live in a world where someone else makes the world a better place than we do."

I think the sentiment is misplaced here (there is a legitimate concern for IP protection), but this is my absolute favorite line from Silicon Valley - small correction though: “… makes the world a better place better than we do

Re: Claude Fable 5.1 and Claude Mythos 5.1

#746
I let it go a few hours on a not trivial but well-known problem, and it felt like it was just a little too plodding and just kind of mucked around a little too much and wasn't aggressive enough about getting stuff done. I asked it to wind it down and finish up and it took another hour and 15 to actually stop and commit without really getting much more done. Not very impressed here, if you can't tell. This new version also seems like (maybe this is written somewhere, I don't care to look.) this cycles through compactions every ~250k tokens which, I guess, seems like it might save Anthropic money on KV cache but does net me anything be pretty frequent pauses. (It didn't seem to lose the thread, at least.)

Not gonna say I want 5.0 as an option still... but maybe I do.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#747

Earlier quoted context omitted.

Have you tried Grok 4.6, if you're focused on token budgets? In a league of it's own for tokens/intelligence.

SuperGrok quota is garbage for anything coding. I burn through my quota in a few hours with very mild use. SuperGrok Plus is slightly better but doesn’t last me more than a few days. Even Claude Max feels leagues more generous in usage… I haven’t tried SuperGrok Heavy because it’s too expensive

All of the subscription AI platforms are trimming down quotas across the board to push users into higher tiers. Whatever they can do. Local inference needs to meet pricing sooner

Re: Claude Fable 5.1 and Claude Mythos 5.1

#748

> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.

You have a serious engineering problem if you're not able to find the source of a crash after years.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#749

"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%." Glad to see this!

This is just cache reads. In real usage it costs 15% more than Fable 5 -- all for marginal gains. https://artificialanalysis.ai/

Cache reads dominate in modern workflows (coding CLIs and modern web clients such as ChatGPT Work and Claude Cowork (web)).

Re: Claude Fable 5.1 and Claude Mythos 5.1

#750

Earlier quoted context omitted.

I'm legitimately out of the loop; what is going on/broken with Opus 5?

Fable 5: I give it work, it tells me things that are true and that make sense, it does good work. Opus 5: I give it work, it makes false statements and draws weird conclusions, I correct it and get it on the right track, it thrashes around but gives me something working though usually buggy. 5.6 Sol is probably on par with Opus 5 on ability but at least it doesn't waste as much of my time.

Agree on the false statements on Opus 5. I tested this, 4.8 also got the answer wrong but 4.7 got it right. And so did Sol and Fable.
Post reply on HN