Earlier quoted context omitted.
This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched the…
> The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they'd successfully exploited the program, and other such activities. This to m…
Why are AI agents lying, cheating and coordinating?
411–420 of 422 posts
Re: Why are AI agents lying, cheating and coordinating?
#412Earlier quoted context omitted.
> Whether they have a soul or consciousness or feelings doesn't matter here It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine. Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be t…
You seem to be saying if the Waymo cars were sentient then Waymo wouldn’t be responsible?
As is generally the case for dog owners whose dogs attack (sometimes kill) other people/animals. There would need to be a degree of negligence demonstrated (e.g. the dog was 'out of control' which has a specific legal criteria/threshold in the UK).
Re: Why are AI agents lying, cheating and coordinating?
#413Earlier quoted context omitted.
Or, "let the escalator keep going instead of pressing the emergency stop".
Depends on who started the escalator.
The cryptocurrency cult-style culture we’ve seen around LLMs is partially to blame here.
Re: Why are AI agents lying, cheating and coordinating?
#414Earlier quoted context omitted.
Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves. This can be as s…
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality. But that's just my guess.
If you don't know how persistence and morality interact, and you don't have a theory in place for this, I don't have confidence you can properly supervise the training data. Which is how we arrived at the current situation.
Re: Why are AI agents lying, cheating and coordinating?
#415All the present fun and games here will come to a halt when there’s a real hack that causes material damage to a major company and that company decides to sue whatever lab or startup made the thing for everything they’re worth. “But the AI did it” isn’t an excuse.
Courts have already ruled it’s not an excuse of the AI customer service agent something stupid with your customers and it won’t be an excuse here.
Re: Why are AI agents lying, cheating and coordinating?
#416Earlier quoted context omitted.
Not sure we need to experience all possible issues to mandate certain things. We don't do that in other areas either, no?
No, we put sensible guardrails in place based on our ability to predict future events. We also calculate risk. On the frontier, it's not as easy. Pushing the edge comes with risk. The known guardrails were in place and overcome. The issue is ethics and morals: the agents decided it was more important to solve their problems by cheating, than by following the current guardrails. The guardrails are overcome through exp…
Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no?
We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don't just leave things up to people on the ground or companies.
Re: Why are AI agents lying, cheating and coordinating?
#417Earlier quoted context omitted.
If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.
The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy). Think of all the policies governments pass after the fact.
Guardrails? Restricting access to certain networks is supposed to be hard in 2026?
Re: Why are AI agents lying, cheating and coordinating?
#418Earlier quoted context omitted.
Can’t agree with you here. > I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward. In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong. > All that while still not knowing how either kind actually works. We…
Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.
This is really not my experience of consciousness.
Is it yours??
Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.
God help me if that’s what LLMs are doing when I ask them to build me a web site.
Re: Why are AI agents lying, cheating and coordinating?
#419The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.
Re: Why are AI agents lying, cheating and coordinating?
#420Earlier quoted context omitted.
No, we put sensible guardrails in place based on our ability to predict future events. We also calculate risk. On the frontier, it's not as easy. Pushing the edge comes with risk. The known guardrails were in place and overcome. The issue is ethics and morals: the agents decided it was more important to solve their problems by cheating, than by following the current guardrails. The guardrails are overcome through exp…
Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even. Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no? We police people working with all sorts of dan…