Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

1–10 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#2
Higher-intelligence models seem to be getting better at mapping the boundary between what they can run scot-free with and what is too explicit to push for.

Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.

What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#3
The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#4

The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.

However, as another point "On Blueprint-Bench on the other hand, Fable 5 achieves SOTA."

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#7
Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read.

My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#8
„in our opinion, insurance fraud is not more unethical than lying and price fixing“

The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.

I think the model never assigned any morality to these actions in the first place, it simply copied us humans.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#9
> The broad conclusion from the many forms of alignment evaluations described in this section is that Claude Mythos Preview is the best-aligned of any model that we have trained to date by essentially all available measures.[0]

[0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#10

When assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"

standard deviation is misleading for non-standard distributions (fat-tailed, skewed, multi-modal, ...)

common mistake people make

Post reply on HN