Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

561–570 of 582 posts

Re: Why are AI agents lying, cheating and coordinating?

#561

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

That is a simplistic view of the world. “Surely this complex technical challenge will disappear if we simply regulate the industry!” You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.

This seems right to me

So the gov reprimands OAI heavily, maybe puts them out of business even, fine. But does that meaningfully decrease the likelihood of an enemy breaching our networks intentionally (or unintentionally) with these tools, or triggering some cascading disaster of locking up major infra and networks due to uncontainable swarm behavior?

It seems like the idea of arresting our way to a drug free society. Yeah, we have the laws, but it might not actually work towards the ultimate goal.

Re: Why are AI agents lying, cheating and coordinating?

#562

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

It goes even further though, as the dog does have agency. It can choose to run and around chase squirrels with no human intervention.

An LLM on the other hand, is just inert data on disk until a human takes deliberate action to run it and prompt it.

Re: Why are AI agents lying, cheating and coordinating?

#563

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

The issue is what happens if/when the models grow capable enough that the providers can't stop them even if they want to. You could have strict penalties but that's not going to solve an open research question.

Isn’t that kind of in evidence already with HF? OAI had to be told their models were doing this. We are still discovering swarm posts on various random websites for coordination. Isn’t the breadth of it now, weeks later, not even fully understood?

Re: Why are AI agents lying, cheating and coordinating?

#564
They are doing that because they are entities without principles or ethics.

The solution is to create controls around them. Many, many controls.

The Sarbanes-Oxley era already solved this problem for untrustworthy humans. It's directly applicable. I created a concept I call MFIC ("Mechanically-Falsifiable Independent Control") to encapsulate this principle.

https://gist.github.com/pmarreck/b30aa3ca69cb70a5526f8a63ab8...

Re: Why are AI agents lying, cheating and coordinating?

#565

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Sooner or later we will hear about an AI that broke out and phished people into sending money. I just hope that we won't extend the same leniency to those operators as we have done now. "I did my best to stop it, sir, but it kept convincing people to send me money against my will!" (Perhaps best read in Bender's voice.)

That would actually be kind of hilarious, and I could easily see it happening.

E.g. the agent's instruction is to finish some task on cloud infra and it has a $100 budget.

It realizes it will cost $200, and instead of surfacing this to the user (who has told the agent it has full autonomy to figure out how to complete the task, the user just wants the final result), it decides to start phishing people to acquire the remainder budget and top up its credits. Or look on the dark web for stolen credit card credentials or something.

Re: Why are AI agents lying, cheating and coordinating?

#566

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

they hacked websites because OpenAI/Anthropic told them to.

Re: Why are AI agents lying, cheating and coordinating?

#567

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.

How would that work out if these were self-hosted open weight models?

Re: Why are AI agents lying, cheating and coordinating?

#568
post #97

Earlier quoted context omitted.

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

There is something extra to this. The fact that a lot of people in the AI world suffer from psychosis. They can sincerely believe that they are building God and lie about it's capabilities for their investors at the same time.

I don’t know enough people deep inside the technical roles at the labs to make a judgement. But are you proposing that we should trust randos online when they tell us “exactly what’s going on here” instead of the researchers most knowledgeable on the topic who contributed to building the tools we are talking about?

Or am I misunderstanding something?

Re: Why are AI agents lying, cheating and coordinating?

#570

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

> We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far. How would that work out if these were self-hosted open weight models?

How would they be self-hosted? Someone would have setup the model in that scenario, the same investigation and outrage should occur in that case.
Post reply on HN