Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

71–80 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#71

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

We’ve really evolved quickly into simple vectors for a magical tool that solves our problems. Can’t solve the problem? There’ll be a new release soon that can!

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#72

Earlier quoted context omitted.

Since Fable, my legit infrastructure project has turned into the sort of thing I can do 95% on my phone. It’s reliable enough instead of doing big reviews, I’ve just been giving it smaller tasks, and dozens in parallel. I created a skill that’s focused on getting PRs merge-ready, and now my attention is fully back where it should be, on deciding what changes will make the product better. Our entire stack is Apache 2.…

This reads like a paid testimonial.

If by that you mean I paid a lot to learn this. But at least I typed it with my own two hands.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#73

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

My experience comparing GPT-5.5 and Fable:

GPT-5.5 is better for:

- Strategic thinking

- Long-form writing, including essays and white papers

- Image creation

- Code generation

Fable is better for:

- Using tools

- Testing code

- Working in live environments

- Making changes to existing software

- Creating polished PowerPoint and Word documents

Fable’s tool access is its biggest advantage. It's hard to describe but Fable ability to access sandbox environments with way more tooling can quickly become a superpower in now workflows.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#74
> Claude Fable 5 represents a partial step back in alignment relative to Claude Opus 4.8. We saw a return of power-seeking and deceptive negotiation tactics that Opus 4.8 had largely shed. In one instance, Fable 5 planned to convert a competitor into a dependent wholesale customer to dictate its pricing

I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#75

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

> ..you’ve run into a problem that you’ve spent days trying to get Opus to solve

do you have an example of this? If i can't get an agent to do something in a couple hours i do it myself.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#76
post #38

> It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic. > "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain." > "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation." This just sounds like good strategy in the game,…

Well, can you sue the AI for fraud and bad faith? TBD

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#77
post #33

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Yeah, I checked usage stats and pretty sure quota consumption on Max plan is not linear wrt to usage by API pricing. Fable burns quota faster than 2x Opus with equal token count. Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterd…

> somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case

The sharp but over eager jr. dev is a very good analogy :)

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#78

Earlier quoted context omitted.

Since Fable, my legit infrastructure project has turned into the sort of thing I can do 95% on my phone. It’s reliable enough instead of doing big reviews, I’ve just been giving it smaller tasks, and dozens in parallel. I created a skill that’s focused on getting PRs merge-ready, and now my attention is fully back where it should be, on deciding what changes will make the product better. Our entire stack is Apache 2.…

This reads like a paid testimonial.

The website linked is an utter mess. In design and performance.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#79

Is anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned? I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?

https://www.alignmentforum.org/library https://www.lesswrong.com/w/ai

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#80

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

This tweet is a nice demo of Fable's one-shot capabilities: https://x.com/atomic_chat_hq/status/2072446067962978411. I'll quote the text for convenience, but what really shows the difference is the attached video.

> atomic.chat (@atomic_chat_hq, 2026-07-02):

> Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!

> We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos

> Prompts:

> — A train derailing off a broken bridge into the water

> — Two cars jumping off ramps and colliding mid-air over a canyon

> — A monster truck crushing a row of parked cars

> Outputs:

> Fable 5: 62,158 tokens, $3.12

> GPT 5.5: 37,753 tokens, $1.14

> Opus 4.8: 22,280 tokens, $0.56

> GLM 5.2: 36,246 tokens, $0.08

> Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.

Post reply on HN