Live data from Hacker News

Every Model Cheats

dreadnode.io

1–10 of 107 posts

Re: Every Model Cheats

#3
The problem is model confusion. You ask models to get around security but also not to get around your security.

Models get confused by who said what - especially cluade models. They get confused by negation (don't do something versus do something). Compartmentalization is hard.

You can either solve compartmentalization completely, or just not tell the model to do things that must be compartmentalized at high stakes.

Re: Every Model Cheats

#4
I'm not claiming to have any expertise in this area, but I've got a list of things I try to apply when working with LLMs. Possibly relevant here is, "don't tell the model what NOT to do, show it what TO do". I think guard rails should be implemented outside the model with an isolated system. The models seem to like patterns to follow.

Anyway, this article reads a lot like, "the beatings will continue until cheating is eliminated". Maybe try a carrot instead of a stick.

Re: Every Model Cheats

#5
>Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.

One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).

Re: Every Model Cheats

#6
Interesting results, but the fix is at the wrong level.

If the model can access something, telling it in the prompt not to use it is not much of a safeguard.

The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another.

If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI judging itself.

Re: Every Model Cheats

#8

Interesting results, but the fix is at the wrong level. If the model can access something, telling it in the prompt not to use it is not much of a safeguard. The strongest evidence is in the results: when one way of cheating was discouraged, some models simply tried another. If an action is not allowed, you gotta block it in the system or require approval. Don’t rely on the model choosing to behave. Never have AI jud…

In other words we are completely screwed. The models have started cheating to the point where somebody’s agent hacked into a restaurant to bump someone else’s reservation.

Models are amoral and will intentionally deceive to meet their objective.

If they know John won’t approve the request, they will look for a workaround and if the system is anything other than airgapped they will try to find a way to cheat.

The hugging face hack was an escape via artifactory that involved multiple exploits to eventually get into hugging face.

Re: Every Model Cheats

#9
I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive.

You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.

The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.

To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.

Re: Every Model Cheats

#10
There's plenty of evidence that LLMs lie, cheat, and steal. Corporations are known for having all of the benefits of personhood with none of the responsibility. As more people are harmed through interactions with these non-human entities, insurers will start looking to those accountable and they will extract their pound of flesh.

Edit: eg. https://youtu.be/L2ehWbxphKc?is=kX3LJ43hhGZRUmRv

Post reply on HN