Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

51–60 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#52
post #24
post #21

Earlier quoted context omitted.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codeb…

> Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...

Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work is getting done by very narrowly focused batches of subagents, with the main session just coordinating and keeping the doc updated.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#53

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

I felt similarly but after using Fable heavily over the weekend and then flipping back to Opus I can feel a difference. Fable just gets more right the first time, guesses right the first time, and follows through better than Opus. Put simply, I could "trust" it more.

Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#54
post #21

Earlier quoted context omitted.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

15+YOE. Fable 5 is well above the level of Opus. I have used it alongside Opus for a range of hard problems, including porting a large static analysis tool to Rust, building various tooling around .pptx and .xlsx documents.

In all cases, Fable clearly outperformed Opus.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#55
I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t.

GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#56

It's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.

It really, truly is. No matter how many trillion parameters it's built on, it's still just a probability model. It's just on a constant loop of guessing the next word with some inputs from a deterministic controller. Any claims of "motive" or "behavior" are inappropriate anthropomorphizing of something that will never be more than a mathematical model of things humans do. It "chose" the corresponding words to describe a dishonest trade strategy based entirely on configured temperature and a series of clock times on the computer running the LLM.

There's probably some quantifiable component of moral alignment embedded in the idiosyncrasies of the English language itself, if one were to dig deep enough, but that's the stuff of MIT doctoral theses and squarely beyond anything most of us is remotely qualified to talk about.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#57
post #24

Earlier quoted context omitted.

I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codeb…

> Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ... Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work i…

For optimizations or proofs I suppose? Wouldn't know why else you would do something like that.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#58

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

Funny how the two top comments are contradictory. We need better than anecdotes to understand what the new models bring.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#59
Is anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned?

I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#60

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

Funny how the two top comments are contradictory. We need better than anecdotes to understand what the new models bring.

Here’s one difference I have seen. I forgot I had a multi-session audio probe running while trying to repro audio glitches, and Fable came back with: “your pops are already on tape.”

Interesting choice of words. Phrased so casually. It picked a low-tech idiom that fit the situation instead of giving some sterile technical answer. That kind of language and context awareness never happened for me with Opus, or gpt 5.5.

Post reply on HN