and therefore any assertions _AT ALL_ about alignment are null and void.
Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
51–60 of 146 posts
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#52Earlier quoted context omitted.
Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.
I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codeb…
Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work is getting done by very narrowly focused batches of subagents, with the main session just coordinating and keeping the doc updated.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#53Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#54Earlier quoted context omitted.
Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.
Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.
In all cases, Fable clearly outperformed Opus.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#55GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#56It's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.
There's probably some quantifiable component of moral alignment embedded in the idiosyncrasies of the English language itself, if one were to dig deep enough, but that's the stuff of MIT doctoral theses and squarely beyond anything most of us is remotely qualified to talk about.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#57Earlier quoted context omitted.
I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codeb…
> Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ... Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work i…
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#58I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#59I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#60I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
Funny how the two top comments are contradictory. We need better than anecdotes to understand what the new models bring.
Interesting choice of words. Phrased so casually. It picked a low-tech idiom that fit the situation instead of giving some sterile technical answer. That kind of language and context awareness never happened for me with Opus, or gpt 5.5.