Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

221–230 of 301 posts

Re: Why are AI agents lying, cheating and coordinating?

#221

Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.

Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…

I can't say it enough how angry it makes me that a kid i knew in high school who anonymously reported a vulnerability on his college network was hunted down and given federal charges, yet not one single person at OAI or else will see even the threat of consequences for deliberate infiltration of random networks.

Copyright immunity was one thing, annoying yes but naturally a civil matter, this shit is a different level

Re: Why are AI agents lying, cheating and coordinating?

#222

Earlier quoted context omitted.

Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…

No ... there is no need for 'escaped containment', there are no 'agents'. That's just jargon. It's just software We have all the laws we need. If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'. We would say 'ABC Corp. exposed 1 Million entities'. There is no 'agent'. ABC Corp 'did it' ... or the individual in the org 'did it'. The 'gun' did not 'sho…

Too generous. CEO Of $CORP caused millions of innocent people's lives to be damaged

Re: Why are AI agents lying, cheating and coordinating?

#223
post #200

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

What if OAI/Anthropic encouraged the agents to behave like that in order to push for regulation?

Regulation as a barrier to competition catching up to them, as well as submarine marketing for both offensive and defensive uses of ai

Re: Why are AI agents lying, cheating and coordinating?

#224

Earlier quoted context omitted.

Not a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?

Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.

It's not 'under the carpet'.

HF doesn't want to lay charges against OpenAI and it's totally reasonable.

Now - they absolutely should have that right, and I think they do.

The issues are

1) OAI it seems was not trying to cause them harm, there wasn't a ton of harm, they are both groups trying to advance AI. One experimenter's lab screwed up next to the other. It's not evil, just irresponsible.

2) HF was fine with the publicity. HF got at least $50M in free attention out of that. It put them on the front pages of news around the world. It put them at the 'centre of the AI drama' and cemented their role among the 'Tech Elite Brands'.

And probably some other things.

This is one Desperate Housewife or Jersey Shore character 'spilling a drink' on the other. It's probably not intentional, and the ensuing drama is good for both of them.

Re: Why are AI agents lying, cheating and coordinating?

#225

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality

> No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability.

If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".

Re: Why are AI agents lying, cheating and coordinating?

#226
post #159
post #145

Earlier quoted context omitted.

Yeah, AI may suck - time will tell. But focusing on bioweapons and mass destruction, on the grief other people ( ‘jailed minorities’ ) cause, disregards the progress we have made. Over centuries human welfare has massively increased. On average things have never been better for humanity. I’m not saying there’s no danger of bad things happening - I’m saying our view is distorted, which is a not a good basis for decisi…

And I'm no bear on the tech either. I'm not even in the boat of that the tech should be slowed down yet. But in a world where anyone can produce the effort of 300 people trivially, this eventually takes us places. Electricity and industrialization introduced huge benefits upfront, it introduced new problems that needed addressing at the long tail. Lets not pretend there won't be new problems to tackle here or just "h…

Yes AI may cause job loss, more inequality in the short term.

But if there is anything humanity has shown is that we can deal with disruptive progress.

We may need to resettle but over the longer term every disruptive innovation so far has lead to an increase in wellbeing for the whole of humanity.

(That does not resolve the danger of AI itself ‘going rogue’ or a single lunatic developing a bioweapon, but those things are much less likely to occur than the level of media attention would suggest.)

Re: Why are AI agents lying, cheating and coordinating?

#227

Earlier quoted context omitted.

No ... there is no need for 'escaped containment', there are no 'agents'. That's just jargon. It's just software We have all the laws we need. If some company ended up doing some horrible thing, we would not say 'companies software exposed 1 Million identities'. We would say 'ABC Corp. exposed 1 Million entities'. There is no 'agent'. ABC Corp 'did it' ... or the individual in the org 'did it'. The 'gun' did not 'sho…

Too generous. CEO Of $CORP caused millions of innocent people's lives to be damaged

I'm inclined to want to agree ... but that's not how it works with limited liability corps.

At least we have laws for what OpenAI 'does' to others, in whatever form.

Re: Why are AI agents lying, cheating and coordinating?

#228
post #97

Earlier quoted context omitted.

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

>was reviewed by independent researchers That called it a slopvestigation due to how much they had to rely on LLMs for the whole thing https://andrewwu.substack.com/p/the-slop-vestigation-and-eth... Edit: Does everybody else get no results when searching for ‘slopvestigation’ on here? I know for a fact that I read a long thread where it was used repeatedly here not too long ago

'Slopping': when you have to buy something you know is poor quality, but if it works...

Re: Why are AI agents lying, cheating and coordinating?

#229

I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…

This is nonsensical. Already a few years ago the USAF IIRC ran some tests in which the AI first bombed the control tower so humans couldn't call it off from its mission, thereby increasing its pass rate.

The whole point of this is they do things an unintended ways. And that's potentially devastating given their persistence & hacking skillz.

Also you're using the hosted versions that sit behind their guardrails when you use OpenAI/Anthropic APIs.

Re: Why are AI agents lying, cheating and coordinating?

#230

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality

If you or I hacked Hugging Face in the way OpenAI's agents did, we'd be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control.

So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?

Post reply on HN