Every Model Cheats
1–10 of 107 posts
Re: Every Model Cheats
#2Re: Every Model Cheats
#3Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.
You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.
Re: Every Model Cheats
#4Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.
Re: Every Model Cheats
#5One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).
Re: Every Model Cheats
#6If the model can access something, telling it in the prompt not to use it is not much of a safeguard.
The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.
If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.
Re: Every Model Cheats
#7Re: Every Model Cheats
#8Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…
Models are amoral and will intentionally deceive to meet their objective.
If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.
The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.
Re: Every Model Cheats
#9You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.