I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
71–80 of 146 posts
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#72Earlier quoted context omitted.
Since Fable, my legit infrastructure project has turned into the sort of thing I can do 95% on my phone. It’s reliable enough instead of doing big reviews, I’ve just been giving it smaller tasks, and dozens in parallel. I created a skill that’s focused on getting PRs merge-ready, and now my attention is fully back where it should be, on deciding what changes will make the product better. Our entire stack is Apache 2.…
This reads like a paid testimonial.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#73I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
GPT-5.5 is better for:
- Strategic thinking
- Long-form writing, including essays and white papers
- Image creation
- Code generation
Fable is better for:
- Using tools
- Testing code
- Working in live environments
- Making changes to existing software
- Creating polished PowerPoint and Word documents
Fable’s tool access is its biggest advantage. It's hard to describe but Fable ability to access sandbox environments with way more tooling can quickly become a superpower in now workflows.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#74I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#75I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
do you have an example of this? If i can't get an agent to do something in a couple hours i do it myself.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#76> It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic. > "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain." > "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation." This just sounds like good strategy in the game,…
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#77Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
Yeah, I checked usage stats and pretty sure quota consumption on Max plan is not linear wrt to usage by API pricing. Fable burns quota faster than 2x Opus with equal token count. Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterd…
The sharp but over eager jr. dev is a very good analogy :)
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#78Earlier quoted context omitted.
Since Fable, my legit infrastructure project has turned into the sort of thing I can do 95% on my phone. It’s reliable enough instead of doing big reviews, I’ve just been giving it smaller tasks, and dozens in parallel. I created a skill that’s focused on getting PRs merge-ready, and now my attention is fully back where it should be, on deciding what changes will make the product better. Our entire stack is Apache 2.…
This reads like a paid testimonial.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#79Is anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned? I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#80Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
> atomic.chat (@atomic_chat_hq, 2026-07-02):
> Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!
> We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
> Prompts:
> — A train derailing off a broken bridge into the water
> — Two cars jumping off ramps and colliding mid-air over a canyon
> — A monster truck crushing a row of parked cars
> Outputs:
> Fable 5: 62,158 tokens, $3.12
> GPT 5.5: 37,753 tokens, $1.14
> Opus 4.8: 22,280 tokens, $0.56
> GLM 5.2: 36,246 tokens, $0.08
> Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.