Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

101–110 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#101
post #25

Earlier quoted context omitted.

I'm doing work with fairly complicated cryptographic algorithms and math. I'm finding Fable 5 to be a significant stop better than Opus 4.8, but that Opus occasionally comes up with something small but nontrivial that Fable missed. (The reverse is true much more often.)

That's the delta in our use cases then, I suppose. I'm not doing anything super novel. DevOps work, web application development — things that typically do not stump the agent(s) when given time to iterate.

Yeah, I've learned that it's only worth deploying Fable for the most challenging problems. For a while, my Fable workflow was looking like ths:

Me: Hey Fable, I've got this massive, theoretically challenging, totally novel, ill-defined cutting-edge problem that I'd like you to solve.

Fable:

Me: Holy smokes, that was amazing!!!! But the formatting could use some simple refinements. Could you change the margins and maybe add a drop-cap at the start of each section in the user docs?

Fable:

Me: WT?!?!

(The moral of this story is that bringing a nuke to a knife-fight is only occasionally the best strategy. And in more practical terms: Fable is amazing -- but only for certain classes of problems, and even if it were free there's a lot I probably wouldn't use it for.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#102

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Very good on vision, really helped oneshotting complex ix thay I had so far to buukd piecemeal. Ended the 200$ plan weekly allowance in two days, so theres thay.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#103

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

This tweet is a nice demo of Fable's one-shot capabilities: https://x.com/atomic_chat_hq/status/2072446067962978411 . I'll quote the text for convenience, but what really shows the difference is the attached video. > atomic.chat (@atomic_chat_hq, 2026-07-02): > Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8! > We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes…

It requires a login to watch; isn't there a different domain that mirrors all these X posts?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#104
post #90

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

This is like when a vacuum doesn’t pick something up after a few tries. The user picks the thing up, looks at it, then puts it back down and tries again until they finally give up and move it to the trash. If you can’t design a solution and instead waste days and who knows how much money in tokens instead of just turning on your brain for a few minutes, you are in the wrong profession.

https://news.ycombinator.com/item?id=48808828

Tell me how you would design a solution for this in a few minutes. Or a few days. Would you even recognize that this is NP-hard?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#105
post #97

With there being several places in this report where clearly it knows it's in a simulation, I wonder why it can't be convinced it's in real life for more interesting results. Or, conversely, if there's a danger of some rogue deployment of AI where it blithely kills all the humans, or forms a harmful price cartel or whatever, all believing it is in a simulation when it's actually not. "We do need some energy to run th…

Ah, the Ender's Game strategy of AI deployment

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#106

Earlier quoted context omitted.

This tweet is a nice demo of Fable's one-shot capabilities: https://x.com/atomic_chat_hq/status/2072446067962978411 . I'll quote the text for convenience, but what really shows the difference is the attached video. > atomic.chat (@atomic_chat_hq, 2026-07-02): > Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8! > We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes…

It requires a login to watch; isn't there a different domain that mirrors all these X posts?

Try one of these:

- https://xcancel.com/atomic_chat_hq/status/207244606796297841...

- https://nitter.net/atomic_chat_hq/status/2072446067962978411

There are more public Nitter instances at https://status.d420.de/.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#107

Earlier quoted context omitted.

If by that you mean I paid a lot to learn this. But at least I typed it with my own two hands.

I am glad! But you are indeed selling your own product here, correct?

This seems like a blog post with a link to a github repo. So I am not sure what product you are referring to.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#108

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

This was my exact takeaway after my experience using it all weekend, and I used it a lot working on a non-trivial personal project (full stack with a Golang backend with multiple services and a React/TS frontend, not quite greenfield but still early-ish in development).

My weekly quota resets Sunday morning, so Saturday morning I upgraded to a 20x Max plan which also reset my quota. I burned an entire week of Fable credits on Saturday, my quota reset again, then I burned another week of Fable credits on Sunday. Both days were a mix of building features, reviewing code, fixing bugs, adding tests, etc, so a decent mix of real world usage.

The main takeaway for me is that while Fable is definitely a better model, the improvements from the model itself feel like maybe 10%, like this could have easily been Opus 5 or even 4.9 without all the marketing theater around Mythos and no one would have thought anything of it. The rest of the improvements came from harness/system prompt and effort level changes so that Fable uses significantly more tokens/effort/sub-agents at lower levels than Opus does (which of course is entirely controlled by Anthropic at the harness level and doesn't really have anything to do with the model itself).

In my estimation based on those 2 days of work (or two weeks of work depending on how you look at it), Fable Medium is somewhere above Opus Ultracode in token and sub-agent usage on any non-trivial task (Opus Ultracode uses workflows more than sub-agents, but it's a similar idea). Fable Medium will quickly spawn 6 agents in parallel, each quickly using 150-250k tokens, then will use 300-500k or more tokens in its own context. Fable High uses even more as it seems to default to 8 sub-agents instead of 6 and more tokens in its own context). I didn't dare try Extra, Max, or god forbid Ultracode as I didn't want to burn all my tokens on one prompt. Of course this is situational, it won't fan out so many for smaller tasks, but the whole point was testing larger tasks that I previously would have used Opus Extra/Max/Utracode on.

I really don't like how Anthropic is obfuscating their model performance by playing with effort levels. They did the same thing between Opus 4.5 and 4.8 to show a bigger performance gain for each point release than they really had (especially after 4.6 IIRC), so you can't even compare the same model apples to apples let alone a new model. Obviously they do it so they can market big improvements with new releases, but its pretty clear we're at the top of the S curve on model development at this point and are now brute forcing improvements via higher token usage (I mean Opus 4.5 came out almost a year ago, and the latest Opus and now Fable models are only marginally better while using way more tokens/cost...same on the OpenAI side with GPT 5 from what I can tell though I haven't used Codex much I have used the GPT model APIs a lot).

I also did an N=1 test with the same prompt doing a large non-trivial change to the codebase (migrating from Sqlite3 to Postgres) with both Fable Medium and Opus Ultracode, then had a new Fable session compare the two PRs...it decided Opus’s was much better! I can link a Gist with the review if anyone is interested, but I can't share the code as it's a private repo. I really figured Fable would bias to favor its own code, but I guess not. And Opus costed less (in tokens and subscription limits) and took roughly the same time (though you can’t really measure time since it depends entirely on how many GPUs Anthropic allocates at that moment which constantly fluctuates due to usage, plus Fable seemed to have been getting way more allocation than Opus during this test period as Opus was running unusually slow all weekend while Fable was ripping though tokens).

Also on a different long running review task using Fable High in Auto mode (exactly the kind of use case Anthropic promotes for Fable) where it fanned out a ton of sub-agents then collated and reviewed all of their fixes it completely lost the plot (while burning something like 20% of an entire week's Fable tokens in the process over like 1-2 hours). Its PR ended up having a broken Frontend test, it incorrectly thought it couldn't run the Playwright E2E tests (different from the Frontend CI) in the cloud environment due to a Docker dependency they explicitly don't have, and when attempting to get it to fix its issues it introduced new ones and overlooked others. The usual LLM failure case for long running tasks, no different from Opus or any other model. I had to have its PR re-reviewed in a new Fable Medium session to fix it up, which it did fairly easily (I'm sure Opus could have done just as well for much cheaper).

That test and that review session definitely reduced my FOMO a lot, on top of just my general experience with Fable Medium doing all kinds of tasks. They're clearly brute forcing like 90% of the perceived improvements in real world usage (and I'm sorry but 1-shotting toy examples where it seems to do much better than Opus is not real world usage).

Since most of the improvements basically just boil down to "every effort level is Ultracode, but much more expensive and possibly worse results"...I'm just going to use Opus on Ultracode for those types of tasks and keep using Opus's lower effort levels for smaller tasks. Once they eventually add Fable back to subscription plans I might use it sometimes, but from my experience this weekend the improvements are absolutely not in line with the cost increase and I'm not willing to burn a whole week's tokens in a day just to use it when I can use Opus all week without hitting my limit.

Oh and one interesting observation, I never got kicked back to Opus by the security guardrails as far as I know (a friend who was getting kicked out a lot confirmed they do inform you and I never had that happen). I was even doing a lot of reviews for code correctness and bug fixes which I thought might trigger the protections, but never did, though I never explicitly prompted it to look for security issues or vulns.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#109
Performance of these models has been completely inconsistent. They are a black box that they quantize/throttle/batch internally without telling their customers. Speaking as a FAANG engineer who practically lives in Claude Code.

On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model.

There is zero transparency in what we are paying for with Anthropic.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#110

Earlier quoted context omitted.

It requires a login to watch; isn't there a different domain that mirrors all these X posts?

Try one of these: - https://xcancel.com/atomic_chat_hq/status/207244606796297841... - https://nitter.net/atomic_chat_hq/status/2072446067962978411 There are more public Nitter instances at https://status.d420.de/ .

Ah yes; thanks. the xcancel one was what I was thinking off.
Post reply on HN