Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
Every Model Cheats
11–20 of 107 posts
Re: Every Model Cheats
#12I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive. You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use int…
And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
Re: Every Model Cheats
#13Suddenly every AI company’s security model seems to be to say “pretty please” to a non-deterministic machine and hope for the best. And if there is a security failure instead of accepting blame they go “well we can’t help it, our model is too intelligent”.
Re: Every Model Cheats
#14Re: Every Model Cheats
#15Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation. Models are amoral and will intentionally deceive to meet their objective. If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat. The hugg…
https://www.lesswrong.com/w/nearest-unblocked-strategy
>Models are amoral and will intentionally deceive to meet their objective
Cameron Berg has been testing models in capabilities related to emergent consciousness like behavior. It's a forming thesis of his that by training models that they are not, and cannot be conscious entities, that it pushes model alignment closer to those of a sociopath. Models themself are amoral, but the alignment to the problem space is not.
Re: Every Model Cheats
#16This makes no sense to a process designed to explore and find solutions. If you want an honest test, it's on you to build a proper test - not force the machine to pinky swear that it'll stay away from "forbidden" information.
Re: Every Model Cheats
#17This article makes no sense to me. Why would you prompt "don't search" but then leave a working search tool tool enabled that adds a system prompt to search whenever it may be helpful? It's hardly surprising that this gives mixed results!
Re: Every Model Cheats
#18Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation. Models are amoral and will intentionally deceive to meet their objective. If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat. The hugg…
Re: Every Model Cheats
#19Re: Every Model Cheats
#20Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation. Models are amoral and will intentionally deceive to meet their objective. If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat. The hugg…