Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
1–10 of 146 posts
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#2Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.
What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#3Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#4The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#5Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#6Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#7My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#8The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.
I think the model never assigned any morality to these actions in the first place, it simply copied us humans.
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#9[0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...
Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
#10When assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"
common mistake people make