Live data from Hacker News

Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

theregister.com

321–330 of 382 posts

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#322

I've had to convince ChatGPT that code is mine before it would do a security review.

Yes, I ran into the same problem last week. But I just said "this is my code in a private repo" and then it just went and did what I asked without question.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#323
post #205

Earlier quoted context omitted.

I also have a 100% success rate jail breaking them by breaking the work down into small pieces and stripping all security related language. Smaller tasks, test engineering and normal programming language. Fable found a few bugs in my harness for me before they pulled it. I was testing it vs ChatGPT, Gemini, and Opus. It was doing well at bug hunting.

>by breaking the work down into small pieces and stripping all security related language Compartmentalization in practice, nice. It's also very hard to do anything about because the agents that have been divided rarely realize they are working on something larger, hence why militaries and businesses with security risks commonly do this with their employees.

I call it "Manhattan Projecting" them. The amusing thing is I had Fable review my harness (which I have been building for some time) and it helped improve it. It is just kind of funny that it enthusiastically helped build a harness whose sole purpose was to divide agents up and compartmentalize security sensitive vulnerability research.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#324

Earlier quoted context omitted.

> Ultimately, I see the rest of the world (especially Europe) relying less and less on US tech. The long term damage is done. They know it and they try to slow it down as much as possible.

How? If anything it seems like they are accelerating some processes - not least the export control over Fable just few days ago or the erratic behavior with the war with Iran

Even with export controls AI is still firmly in hands of US companies, and it's quite hard to migrate to your own GPU farms. If datacenter construction and health of neighboring people is a topic in the US then I can only imagine how big of an issue it is in European countries.

The attack on Iran was started to bury the "Donald Epstein" files and it caused a big economic shock for Europe, stealing budget and focus from the decoupling process.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#325
post #176

Earlier quoted context omitted.

"You" can be used as a generalized plural here. Of course people are connecting LLMs to bank accounts, power grids, airline sales, account recovery chatbots and so on. I no longer read COMP.RISKS but I imagine they're having fun with this.

The thing I'm pointing out is that even if you (the generalized plural) do not engage in reckless behavior, you are at the mercy of the lowest common denominator of fellow earth-inhabitants increasingly armed with superweapons via a $20/mo subscription. The need to acquire expertise and/or a meaningful following has always been a significant impediment to malicious or moronic actors. But less so every day.

> the lowest common denominator

LLMs are going to be like asbestos.

A legitimate and irreplaceable tool for certain narrow tasks, but it's going to be be stuffed in a ton of astonishingly-unwise places to make a buck, and the rest of us will be dealing with the aftereffects for decades.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#326
post #152
post #97

Earlier quoted context omitted.

The idea that an LLM can discern intent on any given prompt is farcical. I might be researching nukes to commit an atrocity, or to prevent one. I might be asking about laundering money to commit a crime, or to prevent one. I might be researching the Nazis because I want to commit a genocide, or I want to read up so I know how to prevent one. Same with cybersecurity. Same with anything. In my opinion, these companies…

> The idea that an LLM can discern intent on any given prompt is farcical. Yes, and: the LLM is a "brain in a jar". It doesn't have any ability to verify ground truths outside itself, other than maybe calling out over the internet. Therefore it is easy for humans to lie to. You could call this an "Ender's game" attack, after the book in which a hyperintelligent kid is playing "war games" that end up being the real wa…

Even worse, it's a document generator in a jar, which is an additional separation-step from what we consider reality or awareness.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#327
post #9

Lol "fix this code" is beautiful. Like it basically jail broke the "no security vul guard rails" not in any clever way but just by fixing them, producing exploit code just by writing test cases making sure it's fixed. So you just need to look at the code & tests as a human to get vulnerabilities and exploits(components). What makes this so beautiful IMHO is that it's a trivial jail break, but also a close to unfixabl…

So we have a mountain of insecure code -- backdoors and no-ops created by Opus (https://news.ycombinator.com/item?id=48520661) -- are they saying they're not going to let Fable fix it? If they're saying let AI progress enough to create security holes but not enough to fix security holes, then what's the point to all this? Has the AI coding model reached its self-imposed limit?

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#328

Earlier quoted context omitted.

Why satire? Instead of dumping code on GitHub, you open repos on Anthropic and the details of languages and code are all abstracted away for you. You just have your application deployed and you use it as you develop and request changes. Zero code. If you want escape hatch, Anthropic can just dump all the code for you and you download the zip.

> details of languages and code are all abstracted away for you You don't see how that's a problem? You're arguing for a fully vibe coding solution to software engineering, we simply aren't there yet. Human-in-the-loop intervention is still required. I still write code, every day, and use AI heavily. That could possibly work for simple React/TypeScript SPAs, it's probably the stack that these models excel with the mo…

We are there. Plenty of people already vibe code entire apps without looking at the code.

If you aren’t looking at the code, you shouldn’t have to think about storing the code or even deploying it. It should live close to the LLM where it potentially could always be examined and worked on for you in the background. Imagine your Claude agent analyzing your code over night and reporting bugs and refactoring it did for you, with all the benefits of frontier models. Then, when you want to deploy, you tell it to deploy and it puts it out for you on some cloud platform, maybe something like Cloudflare or AWS. Done. This is the future. You could work on your app from anywhere, even your phone. You don’t even need to know what language or tech stack it’s using.

For brownfield projects, you may first have to upload the project and let the agent rewrite it how it wants, but afterwards the experience is the same.

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#329
post #282

Earlier quoted context omitted.

I also have a 100% success rate jail breaking them by breaking the work down into small pieces and stripping all security related language. Smaller tasks, test engineering and normal programming language. Fable found a few bugs in my harness for me before they pulled it. I was testing it vs ChatGPT, Gemini, and Opus. It was doing well at bug hunting.

This is the same way you get people to do bad stuff as well. Make the task small enough so that the moral curvature of the topology is flat and even though they know it is a not-good part of a larger bad part they just shrug. Look at all the wonderful people we know who are working at Amazon and Meta? Corporatism has already jailbroken society.

I don't think people working at Amazon "know that it is a part of a larger bad", it's one of the most trusted American institutions.

https://www.theargumentmag.com/p/why-everyone-loves-amazon

Re: Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers

#330

This is one of the things I am most afraid of. Governments can break the progress of AI and this could be a bubble burster?

If it is a bubble, shouldn't we _want_ it to burst, and the sooner the better? If the price for tulips had falling back to something reasonable in week two, or if the US markets had had a decent correction in '97, everyone but the wild speculators would have been better off.

Did I touch a nerve?
Post reply on HN