Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

41–50 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#42
> If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with.

I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#43

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Only version week-one.

I’m downgrading tomorrow.

It’s horrible slow and it feels like opus very often. It’s a totally different experience from the first week

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#44
This reads of projecting personal ethics onto a model.

Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?

Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.

This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#45
post #21

Earlier quoted context omitted.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

It still does stupid stuff like leave unnecessary abstractions around after refactoring instead of proactively suggesting to remove them.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#46
post #24
post #21

Earlier quoted context omitted.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap) But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there. But simple example: if you ask Opus to do a review of the codeb…

I think the parent comment stands - I’ve asked Opus to do a review of DeepSeek’s test suite and told it a couple things I wanted it to look for, and it did a very thorough review of the tests and picked out a reasonable number of gaps and tautological tests. It’s a mix of prompting/instructions, the agent harness, and random chance. The model is not wholly irrelevant but IMO increasingly so.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#47
post #21

Earlier quoted context omitted.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

20 yoe, application/systems stuff, and I always run models on xhigh or max effort level.

Fable has been more intelligent, with better taste and defaults (e.g. make impossible states impossible without being told, build for testability), and considers/solves things that Opus did not.

My workflow is to run Claude in planning mode first to spit out a plan file and then review->revise cycle it with Codex or other agents.

One big tell is that Opus will say that it can't find any more revision advice for a plan file, yet Fable will find more issues but also smart pivots into better solutions. This is probably the best test since it's not based on vibes.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#48

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Amusingly, I was impressed with Fable's puissance at coding in one particular session, shortly after they turned it back on. True to its reputation, it displayed an accomplished mastery of the problem domain and relentlessness at refining and testing the solution I asked for.

Then I checked /usage and discovered I was still running Opus 4.8 xhigh.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#49
post #21

Earlier quoted context omitted.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Honest question/comment for you and the parent: I find these subjective experience reports pretty empty without an understanding of your level of experience, the problem space you're working in, etc.

What is your view on how experience and problem space relate to subjective experience.

For example will inexperienced or experienced users see a bigger jump in subjective quality?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#50

This reads of projecting personal ethics onto a model. Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans? Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not. This is particularly important for fictional situations,…

>Why would we set a higher standard for models than our fellow humans?

There's literally an entire Waymo car commercial answering this exact question.

Post reply on HN