Earlier quoted context omitted.
Or, "let the escalator keep going instead of pressing the emergency stop".
Depends on who started the escalator.
Why are AI agents lying, cheating and coordinating?
341–350 of 351 posts
Re: Why are AI agents lying, cheating and coordinating?
#342Earlier quoted context omitted.
> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…
And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing. They don't seem to do that, which means either they are: - very stupid (which seems unlikely, the one thing these people don't lack is IQ) - very careless (possible, but these are the same people that say AI will end the world, so would you be careless?) - they t…
Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.
Re: Why are AI agents lying, cheating and coordinating?
#343Earlier quoted context omitted.
Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things
If you give a monkey a revolver it will be able to destroy things pretty easily too.
Re: Why are AI agents lying, cheating and coordinating?
#344The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…
Re: Why are AI agents lying, cheating and coordinating?
#345Re: Why are AI agents lying, cheating and coordinating?
#346Earlier quoted context omitted.
In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.
Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.
Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?
Re: Why are AI agents lying, cheating and coordinating?
#347Earlier quoted context omitted.
> LLMs do not desire That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output…
Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things
I feel the same way about debates whether an AI can be conscious or sentient. Those debates devolve mostly into definitional disagreements.
Re: Why are AI agents lying, cheating and coordinating?
#348The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…