Live data from Hacker News

Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

theregister.com

331–340 of 382 posts

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#331
post #57

If you set aside political menace, this is a huge problem with Anthropic's strategy. You _cannot_ say that Mythos is super dangerous and can only be rolled out to certain people, but then release Fable with anything other than bulletproof cyber denials. Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work. So you've ended up in a situation where Anthropic are simultaneously claiming it's a…

> Clearly with LLMs, bulletproof denials are ~impossible due to the way LLMs work Exactly. AI safety is nonsensical. You cannot define the set of "bad strings". The billion monkeys with typewriters are eventually going to be able to produce them. Any "safety" system for constraining LLM output is going to have a nonzero leak rate. But on the other hand, this is also irrelevant, unless you're irresponsible enough to c…

This whole thing seems nonsensical. If Mythos is this super hacker by far the best thing to do is just release the dang thing. Donate as needed to cover the Curls of the world but otherwise fixing bugs is (usually) trivial once they're found. Maybe we see a bump in zero days but long term the effect is much securer code.

Playing this game where everyone is blocked by a wall with massive holes in it is absurd. A farce level affair. The black hats will grind their way through prompts while the white hats are blocked from doing a "mythos hack my app" prompt and finding their vulnerabilities.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#332

Earlier quoted context omitted.

> details of languages and code are all abstracted away for you You don't see how that's a problem? You're arguing for a fully vibe coding solution to software engineering, we simply aren't there yet. Human-in-the-loop intervention is still required. I still write code, every day, and use AI heavily. That could possibly work for simple React/TypeScript SPAs, it's probably the stack that these models excel with the mo…

We are there. Plenty of people already vibe code entire apps without looking at the code. If you aren’t looking at the code, you shouldn’t have to think about storing the code or even deploying it. It should live close to the LLM where it potentially could always be examined and worked on for you in the background. Imagine your Claude agent analyzing your code over night and reporting bugs and refactoring it did for…

> and let the agent rewrite it how it wants

So let the agent rewrite decades of battle tested hardware integration code and drivers? Something tells me that's not going to work out right.

Tell me you only make webapps without telling me you make webapps.

I use these models every day in my job. Trust me, we are definitively not there for anything more complex than an React SaaS project.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#333

I'm not sure I've understood it correctly. So, basically the model didn't agree to expose possible vulnerabilities but agree to patch those? Regardless of the request to take Fable 5 down. Why is requesting the model to show vulnerabilities is being blocked if fixing it not? is it based on the assumption of the intention? I don't quite get the benefit of limiting it. So if anyone can explain it better it'll be apprec…

> Why is requesting the model to show vulnerabilities is being blocked if fixing it not? This is how Anthropic describes Fable's behavior: "When Fable’s classifiers detect a request related to cybersecurity, biology and chemistry, or distillation, the response is automatically handled by Claude Opus 4.8 instead. Users will be informed whenever this occurs." So if you ask the model to "find security issues in this cod…

> I guess the "exploit" here is that if you just tell Fable to "fix this code", which is not "a request related to cybersecurity", it will fix security issues (as it should).

The original sin is calling any bugs security bugs in the first place.

It's just unintended behavior.

If you say "should this model be able to fix unintended behavior" the answers are not alarming.

If you say "what about when those behaviors interact in unforeseen ways, allowing even crazier unintended behavior, should it be allowed to help you fix that too?"

Again, the answers are going to be clear.

Our tools must support correctness and resilience and help the exact thing humans are bad at: combinatorial explosions of subtle lacks of correctness…

…and just f'ing fix it.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#334

Earlier quoted context omitted.

We are there. Plenty of people already vibe code entire apps without looking at the code. If you aren’t looking at the code, you shouldn’t have to think about storing the code or even deploying it. It should live close to the LLM where it potentially could always be examined and worked on for you in the background. Imagine your Claude agent analyzing your code over night and reporting bugs and refactoring it did for…

> and let the agent rewrite it how it wants So let the agent rewrite decades of battle tested hardware integration code and drivers? Something tells me that's not going to work out right. Tell me you only make webapps without telling me you make webapps. I use these models every day in my job. Trust me, we are definitively not there for anything more complex than an React SaaS project.

AI friendly code is more important than battle tested code. The sooner you start the conversion the less behind you will be later.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#335

Meanwhile Deepseek V4 Flash will happily hunt security vulns at almost 0 cost. We are ceding the bug hunting to the open weight models.

Deepseek isn't just open weight. It's open source and they even publish research papers alongside them going in depth about their techniques.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#336

They weren't freaked by anything, it's a retaliatory shakedown after ideological differences and Anthropic not doing exactly what they're told/what the Admin wants them to do.

Yep, people are expanding way too much mental energy on basic bribery. Anthropic will agree to work with the DoD, WH insiders will get some lucrative pre-IPO allocation and Fable will be magically "fixed" and available again.

And until then we’re left with braindead Opus 4.8 where I need tell it 7 times before it does something correctly where Fable 5 just did it in the first prompt.

Example: Hey Opus, I’m dealing with this issue on AD and users experience this thing, I tried these. Opus responses with the most braindead call center style respond I’ve ever heard.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#338

As an European, I really don't get where this strategy wants to take the USA to. It's pretty clear everyone is getting scared about changes like this that happen overnight, without clear reason and completely unpredictable. Business requires a stable environment, and Trump is making everything in his power to disrupt business stability. Ultimately, I see the rest of the world (especially Europe) relying less and less…

> Business requires a stable environment

Someone: “You’ve got some nice stable business there that competes with some of the other companies I happen to …”

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#339
post #88

I haven't been following this story, but the US wanted claude to not be able to find bugs in code?

It basically as if you asked it to find ways to enter someone's house and it refused. But then give it exact copy of their house, ask to secure it, which it does and look at what it secured to find out how to get into the original house.

So I was in their house to make blueprints, then I left it, and now trying to get back in?

Kidding aside, it practically requires an open sourced project to a certain extent. Regardless, having worked with braindead Opus 4.8 again since this event and missing Fable 5 with every response I received.

Feels like Anthropic got a major jump in user base and got knocked out by the friends of the competition.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#340
This comment thread really has me thinking. Is it possible that we might be at peak "consumer AI" in terms of intelligence? If it's basically impossible to verify that security-proficient AI is used for beneficial purposes, then these frontier models might start being regulated like WMD. We end up with two tiers of models. Dumber consumer models that are essentially lobotomized to the point of being completely safe. And actual frontier models that are heavily scrutinized and regulated and treated like nuclear weapons.
Post reply on HN