Live data from Hacker News

Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

andonlabs.com

131–140 of 146 posts

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#131
This particular situation just feels like Fable is able to figure out that it's in a simulation, so it's just playing the game it's been put in.

That's not say I think a machine should ever be in a situation where it is allowed to make ethical decisions with real world results, I don't. At least, not given the current basis of the technology. I mean, generally, I don't see how LLMs can ever be capable of making ethical decisions or trustworthy in that role, no matter how good they get at what they're good at. There would need to be a fundamental change in how AI works for me to change my opinion on this, I think. They are ephemeral, they can never experience consequences, they can never want or need anything, thus there is no mechanism for them to take responsibility for decisions.

Anyway, I think the Andon Labs stuff is kind of a stunt, mostly, and I wish somebody would give me a few million bucks to dick around with the little thinky guys in my computer letting them do silly things.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#132

Is anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned? I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?

The "OG" alignment research that MIRI were publishing long before LLMs burst into the scene spent most of it's time on that question. "How can we even define what an aligned AI should do, if human's are not aligned with each other?" as well as "What does being aligned mean when you're a wizard box who's main influence on the world is to create stronger wizard boxes?" and other deep philosophical questions. They came…

Seems like CEV replaces one problem (“what does humanity want?”) with more problems that are probably even harder to answer.

First of all, calling it “coherent” extrapolated volition presupposes that there is such a thing. It doesn’t actually address the objection above, that there may be no such thing. It’s a bit like saying you solved car safety by presupposing a safe car.

Second, it assumes that such a thing can be effectively measured, and there will be no problems or controversies with the extrapolation process itself. There may be several EVs to choose from, and at that point the framework has nothing to say. Maybe we just pick at random then I suppose.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#133
post #75

I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t. GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

> ..you’ve run into a problem that you’ve spent days trying to get Opus to solve do you have an example of this? If i can't get an agent to do something in a couple hours i do it myself.

Here's one:

I was working on an SDF-based CAD tool. One of the things I want to be able to do is select a pair of surfaces (which are identified by a "surface id" propagated up the expression tree) and add a blend between those two surfaces (e.g. a fillet or chamfer).

Here is a video demo of how far I got by doing it myself and using o3 (I think?) to help: https://www.youtube.com/watch?v=LOvqdlDbkBs

The video is a bit confusing because there was some screen-recording lag so it sometimes looks like I clicked on something other than what I clicked on.

You can see that the strategy I have implemented there works most of the time but at the end it fails to apply the blend.

That strategy is to rewrite the expression tree using distributivity so that blend arguments are siblings, and then apply the blend at the union/intersection (min/max) that is their parent.

But this fails when you need conflicting pairs of blends.

The problem is: given an expression tree describing an SDF, (but where the value passed up the tree is a tuple `(distance, surface_id)` rather than just distance), and given a set of fillets of the form `(surface_id_1, surface_id_2, radius)`, produce a new expression tree which fillets all of the places where those surfaces join.

In ambiguous situations, for example the 2 surfaces come together at an edge, and then that edge runs into a 3rd surface, I don't mind how you resolve the region near the 3rd surface as long as it is intuitive and predictable for the end user.

I spent quite some days working with various agents to come up with a solution to this and still haven't managed to find one.

Maybe you could do it in a couple of hours yourself?

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#134

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

I had a bug that both ChatGPT and Opus 4.8 failed to solve, but Fable solved it quite effortlessly.

Anecdotal, sample size of 1.

The only reason I tried fable was because Opus 4.8 went down the same line of reasoning about it as ChatGPT did. Fable solved it a lot faster than the other 2 spent looking into "false clues".

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#135

> If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with. I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?

Believing and knowing are overlapping sets, imagine what you think of when someone says an AI "knows" something, it's the same mechanism (I'd describe it as something along the lines of "encoded abstractly in the weights")

I thought about this more and realised your question might have been "what's the difference between knowing and learning". IE, how can we say the model believes something without having been taught it.

I think you're right that they're basically the same thing. I'd argue they're very slightly different because what an AI model ends up knowing isn't perfectly predictable based on what they were taught (emergent intelligence), but the sentence you quoted is using believing and learning to mean the same thing, it's just trying to draw attention to the fact that the training process structurally enforces "cheat as much as possible without getting caught".

IE, the contrast in the original sentence wasn't "believe" vs "learn", it was "good" vs "permissible"

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#136

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

> to be fairly unimpressive

I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).

Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#137

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

For coding I'm finding the same thing. It does appear better when I'm doing research. But 4.8 with ultracode is very competent at 99% of tasks I throw at it.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#138

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

This was my exact takeaway after my experience using it all weekend, and I used it a lot working on a non-trivial personal project (full stack with a Golang backend with multiple services and a React/TS frontend, not quite greenfield but still early-ish in development). My weekly quota resets Sunday morning, so Saturday morning I upgraded to a 20x Max plan which also reset my quota. I burned an entire week of Fable c…

> I also did an N=1 test with the same prompt doing a large non-trivial change to the codebase (migrating from Sqlite3 to Postgres) with both Fable Medium and Opus Ultracode, then had a new Fable session compare the two PRs...it decided Opus’s was much better! I can link a Gist with the review if anyone is interested, but I can't share the code as it's a private repo. I really figured Fable would bias to favor its own code, but I guess not. And Opus costed less (in tokens and subscription limits) and took roughly the same time (though you can’t really measure time since it depends entirely on how many GPUs Anthropic allocates at that moment which constantly fluctuates due to usage, plus Fable seemed to have been getting way more allocation than Opus during this test period as Opus was running unusually slow all weekend while Fable was ripping though tokens).

Haha I just gave the exact same prompt to Opus Ultracode and it thought Fable’s was better.

Obviously this isn’t the most scientific test due to LLM non determinism, and I still need to manually review both to make my own decision, but the fact at least they seem to basically be a wash is pretty telling about how much of an improvement Fable is when you actually compare them as close to apples to apples as possible (aka similar actual effort/token spend/sub agent activity)

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#139

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.

> to be fairly unimpressive I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code). Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same…

Tried similar thing yesterday in codex, I got "can't show this message" just when it was about to show me some PoC buffer overflow payload. I told it to just put results in my working directory and everything worked.

Re: Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability

#140
post #22

>power seeking is considered an undesirable trait in the context of a business How do you maximize profit while minimizing power?

The whole point is to not maximize JUST the profit. For normal people, it's not all about money, it's also about the society in general.

I'm doing my part to contribute to the microplastics harvest.
Post reply on HN