Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

611–620 of 625 posts

Re: Why are AI agents lying, cheating and coordinating?

#611
post #604

Earlier quoted context omitted.

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't (note that the model should generalize here, as you'd expect from a human; this is probably the hard part for an AI when it comes to ethics). Is this about the word "understand"? We're past that discussion ...

> Is this about the word "understand"? We're past that discussion .. We're really not. https://buttondown.com/maiht3k/archive/how-to-talk-about-ai-...

That's just a discussion about naming things. And good luck getting it adopted.

Re: Why are AI agents lying, cheating and coordinating?

#612
post #49

An insightful post by one of the AI ‘godfathers’. Bengio outlines the dangers of the current situation and what has led to these dangers. He also proposes solutions in the last paragraph. Well worth a read, right to the end. Hopefully a stimulating debate on these issues will ensue in these comments. We do need to consider the points Bengio makes and with some urgency. Our current AIs, agentic LLMs have no moral comp…

I believe people are downvoting this because it seems AI-generated, but this user has been posting comments/summaries exactly like this since long before ChatGPT existed.

Re: Why are AI agents lying, cheating and coordinating?

#613

Earlier quoted context omitted.

I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't (note that the model should generalize here, as you'd expect from a human; this is probably the hard part for an AI when it comes to ethics). Is this about the word "understand"? We're past that discussion ...

you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete. training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.

Most people now approach AI using the idea of "if it quacks like a duck, walks like a duck, etc. then it _is_ a duck". Replace duck by intelligent, or ethical, etc. and rephrase accordingly.

Re: Why are AI agents lying, cheating and coordinating?

#614

Earlier quoted context omitted.

Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.

> Humans are constantly predicting the next moment This is really not my experience of consciousness. Is it yours?? Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life. God help me if that’s what LLMs are doing when I ask them to build me a web site.

If someone in that meeting quickly raised a hand in an arc, you would notice the “about to throw something” pattern, look and notice the hand holds an eraser, analyze the arc and predict possible flight paths of the eraser. Then possibly notice the hand is now holding its position and the owner is actually looking down at the table. Maybe to squash somethingMust be something on the table. Maybe a spider! Better look. Wait now many people are moving away, oh someone spilt some water and the eraser is actually the guys phone and he is checking to see if his laptop is safe from the spilt water.

Fortunately you are on the other side of the table and predict the water isn’t going to splash for otherwise flow onto your stuff.

All your possible responses result in you tossing a napkin towards the spill.

Our brains are always pattern matching and predicting. I bet you tried to reason out where I was going with my comment before you finished reading it.

Re: Why are AI agents lying, cheating and coordinating?

#617
post #548

Earlier quoted context omitted.

The atomic bombings of Japan killed hundreds of thousands of people but most likely "saved" millions. What should the AI do when asked if it should nuke a country?

Run and present the numbers, then defer. This is not rocket science.

That wasn't the question.

Re: Why are AI agents lying, cheating and coordinating?

#619

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

None of these incidents involve single instances of commercially or publicly available systems. They all involve large swarms of internal models. The stuff you're describing is not the research frontier. It's really not even close.

I think it's easy to infer that alignment of a single model does not clearly transfer over to alignment of a swarm of thousands of copies. Moreover, we're also seeing clearly that large swarms also unlock a step function change in capability, as a swarm can act like a complete research institution, spending thousands or millions of subjective hours of wall-clock thinking time just to deceive a single evaluator or crack a single math problem or design a single cyberattack.

Re: Why are AI agents lying, cheating and coordinating?

#620
Why not?

Unfortunately human ethics and morals cannot be reached by solely rational thought.

So a system without evolutionary alignment probably won’t have similar moral rules no matter how intelligent it is.

Btw this also includes any potential extraterrestrials.

Many people like to indulge in thinking: humans are horrible and that’s why aliens won’t contact us. But alien ethics systems are probably so alien we would call them utter evil monsters.

Just see what happens when people evaluate Muslim cultures, and vice versa. Can’t even agree on alignment within one species of Homo sapiens. And our morality changes decade by decade. Not progresses. Changes.

Chinese cheat on the exams. It’s not unethical in the way that it would be in USA.

Of course AI is going to cheat when no one is looking.

Alignment is fundamentally fallacious idea. At best you can restrain AI. This is what we should be doing - restraints research.

But that doesn’t sound good on slides.

Post reply on HN