Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants. Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security fail…
Every Model Cheats
81–90 of 107 posts
Re: Every Model Cheats
#82Before LLMs we had a pretty good idea of security boundaries in software. Applications didn’t trust user input. Operating systems didn’t trust applications. Services and processes didn’t trust each other. There were always tokens, scopes, delegated grants. Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security fail…
Amen. Why can’t we give agents a shell with permissions for programs and file system access controlled by Unix permissions? This seemed to be a solved problem back in the systems where many users were logged into one machine and the admins had to keep everyone from impacting each other.
Re: Every Model Cheats
#83There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh. Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv
Re: Every Model Cheats
#84I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…
>To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat. I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed. Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected…
Re: Every Model Cheats
#85There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh. Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv
Insurers are corporations. What makes you think they just won't pay out or will cease being useful as anything but value extractors? Honestly, the fetishization of "Insurance will save us" needs to die. The risk doesn't go away.
Re: Every Model Cheats
#86I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…
>To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat. I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed. Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected…
I don't understand. We do tightly control it. We can do this perfectly fine. They could have just not given access to the internet.
I'm not against regulating cars, but it sounds to me this is trying to control the car speed by regulating the oil wells.
We dont even have the framework to propose regulation, and you have to hedge it with "may be needed".
And for the people who'd counter that the existential risk is too high - i don't see it. All those stories go something like: "Caveman Bob invented fire today, and tomorrow he'll stumble on room temperature fusion and lasers; marking the beginning and the end of his rise to global domination - therefor we should stop Bob the moment he discovered fire".
Re: Every Model Cheats
#87I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…
How on Earth can you fail to see the danger of not being able to train any kind of ethical framework into very powerful models? If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do.
This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.
Re: Every Model Cheats
#88Earlier quoted context omitted.
Obviously to you perhaps. I've not seen anything that scares me, except for human idiocy. Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime. The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case. So that side…
>The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case. How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions? If insecure softwa…
So what do you mean "tricked"?
Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.
Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.
Re: Every Model Cheats
#89I think this is a great argument against their "intelligence," and explaining why this happens is a really good way to push against the anthropormophization. They don't "know" things, and it's even fair to say "they don't know how to follow instructions," not in a way that humans do. Spicy auto-complete. If they're working in the realm of "how to break into stuff," they're going to see ALL THE WORDS about breaking in…