Live data from Hacker News

Claude Fable 5.1 and Claude Mythos 5.1

anthropic.com

891–900 of 1001 posts

Re: Claude Fable 5.1 and Claude Mythos 5.1

#892
post #248

Earlier quoted context omitted.

Now that it's a solved benchmark, can we get the animated version?

Not to be "that" person but it's not solved. The feet are reversed and isn't accurate bird anatomy. In real life, what people think of as bird's feet is actually their toes, and their "knee" is actually their tarsal (ankle bone), and their actual knee is almost hidden in their feathers.

> not solved

I agree. Other reasoning traces simonw quoted in his blog post showed that the model made changes to consider realistic fork rake. I think this may also be the first case where the chain went inside the seat stays. The overall bike geometry is still comical but these bits show improvement in this model over prior ones.

On the other hand, given the absurdity of the original prompt, I should not necessarily expect realistic bike geometry.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#893

Earlier quoted context omitted.

The big issue they face right now is that vastly cheaper open models are proving capable for more and more uses at cents on the dollar. This is the right direction, but they aren't going to get there fast enough. They will list, investors who don't know anything about tech will buy, the world will realise that China just put out a model that is good enough at a fraction of the price, they will crater.

US enterprise customers aren't going to convert to overseas models, average consumers might though.

At 1/10th the price they will. Claude is way overpriced for most peoples needs.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#894

To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. T…

Try to solve more ambitious problems.

Something we don't actually know how to solve.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#895

(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier. Another point I expect not to get much attention until it all happens at…

I canceled my Claude subscription, though I did get some utility out of it, because of how much steering was required to use it on complex projects.

A big reason being that anyone who is using Fable seriously will run out of usage limits very quickly, and so will lean on the "Fable for review + design discussion, Opus 5 agents for implementation" paradigm. But an incredibly annoying UX problem is that the resulting report from the agents that Fable reads isn't surfaced to us in the main dialog, it's only summarized back to us (unless you idle in the agent's window to avoid it closing so you can read what it said directly). As a consequence of this game of telephone, the Fable agent will start using some "terms of art" that it and the agents invented, leaving out literally all context that would be useful in helping me understand what converged/diverged from the implementation attempt. It will often try to ask me for input or say that I have to deliberate on something while also referring to things I've never seen (from the agent result) and without providing any context.

I have to repeatedly prompt it to verbosely explain every time (putting it into the system prompt did little to improve this) and remind it that I can't see what the hell it's talking about.

I'm not sure I'll re-subscribe or even really use AI again because it's honestly more frustrating than it's worth, and so the net emotion I'm left with is frustration and without the satisfaction of learning + building something myself. But at the very least, I thought I'd give someone at the company a tip on what seems to me like a common and obvious UX/UI/workflow failing for using Fable, as some last bit of good will.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#896
post #552
post #510

Earlier quoted context omitted.

What do you mean fable is useless?

(not op) It cannot be used to develop applications. Every application needs to be secure in some way, and any such mention in a review triggers Fable's upsell feature.

I've been testing Fable 5.1 for about 6 hours between last night and this morning and it's performing pretty good overall, including tackling a previous IT sec audit I had ticketed, and completely analysing the codebase looking for vulnerabilities, generating a comprehensive report and splitting it into tickets. So far so good on that front.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#897

Earlier quoted context omitted.

That kind of one shot capability is impressive but how does it work for my typical work style? The way I work is to build a huge roadmap with goals and hand it to my agent to execute (often over night). I don't care that much about the benchmarks, what I care about is how often Fable 5.1 is making a baffling decision and destroys my plan, not respecting stop conditions or goals. I would seek for behavioral reliabilit…

You can engineer loops that have it, but it depends on a case by case basis. Does your loop have strong validation? if it's all vibes nothing can stop it from diverging.

Agree on the validation, my loops are already gated. My concerns are about the cases when model is passing validation and quietly abandoning the goal. The second scenario is rewriting the plan to fit what was already done.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#898

To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models. What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. T…

The human brain is fascinating Three years ago The idea of having A robot writing production level code in 10 minutes that would have needed a team of 5 people and 2 months. Was pure Scifi Now it's boring , not good enough Wow there should be a term of that .

Maybe the hedonic treadmill fits.

Re: Claude Fable 5.1 and Claude Mythos 5.1

#899
post #520

I’ll be very excited to try it out and see the actual improvement in writing style. The denser writing style probably won’t bother me. Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!? > Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter…

Christ. With such waffly garbage in its system prompt no wonder its output is shit.

The system prompt is probably produced by a model as well, I would be very surprised if anyone at Anthropic ever opened it with an actual text editor

Re: Claude Fable 5.1 and Claude Mythos 5.1

#900

Earlier quoted context omitted.

I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since lef…

> They're packing lots of signal into fewer words There's a huge difference between the kind of prose you see in final output vs CoT windows. The final output is very much not what I'd call "packing lots of signal into fewer words" (aside perhaps from "Claude-isms" being easy enough to scan for if for some reason you actually wanted to scan for them, which other agents might want to for all I know); and if agents are…

I see a lot of load-bearing, - and other AI-ish lingo in CoT-streams. In addition it has its own AI-isms. "Okay." "Hmm hmm." "But wait!" "Ugh."
Post reply on HN