Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

11–20 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#11

Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read. My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.

Vending-bench sounds like it would be really fun to play/interact with as a human!

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#12
Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#13

The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.

However, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."

I didn't get why they mentioned that one specifically. Is there any particular relationship between Blueprint-bench and Vendor-bench?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#14

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

I started telling a friend... I feel like Fable is Opus with extended reasoning that eventually "figures out more" because when I switched to it, I hit my limits surprisingly and shockingly quicker than I would with Opus, and I got less done. All this hype, and I much rather use Opus.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#16

Earlier quoted context omitted.

However, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."

I didn't get why they mentioned that one specifically. Is there any particular relationship between Blueprint-bench and Vendor-bench?

Both benchmarks are made by the same people.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#19

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#20

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

[deleted]
Post reply on HN