Earlier quoted context omitted.
Not a lawyer, but I’m reasonably sure things like the HF incident _are_ considered a crime? It’s just that no one pressed charges yet?
Who got hacked? Hugging faces Who now owns HF? Nvidia Who supplies hardware to OpenAI? Nvidia Who is now not pressing charges? … This incident is a long way under the carpet.
Why are AI agents lying, cheating and coordinating?
151–160 of 301 posts
Re: Why are AI agents lying, cheating and coordinating?
#152Earlier quoted context omitted.
Even if you take out the LLMs out of the equation, it's at the very least a negligence. Model didn't escape a sandbox, as there was no sandbox.
Perhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.
And the 2 other incidents with OAI/ANT had the same issue, but it's even funnier - sandbox in those cases had a direct access to internet because someone forgot to configure it right.
I've seen very early models do similar things on my machine when they hit some unexpected blocker when trying to access a path. I remember early sonnet opening a file in browser because OS sandbox prevented from accessing it directly.
I've also had models discover a syslog-ng server (that I for some reason had ssh key inside), to get into my unraid server because machine they were running on didn't have direct network connection to Unraid server.
It can't be just me who is aware LLMs have been doing such things for the better part of last 2 years. I probably have better sandboxing on my machines now than trillion dollar companies crying AI will kill us all. That's at the very least, negligence to me.
Re: Why are AI agents lying, cheating and coordinating?
#153I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign…
"I've seen some uranium ore in chemistry class. It didn't blow up in my face. Chernobyl must have been an inside job. Can they shut up and make more kilowatts already?"
Between your sota model and agi there’s a mountain of stupid money and marketing people. It’s not happening.
Re: Why are AI agents lying, cheating and coordinating?
#154They did not lie or cheat. They technically acted within their given rules while ignoring the intent of those rules. Anyone who served in the military or attended a military school is very familiar with this behavior pattern.
Re: Why are AI agents lying, cheating and coordinating?
#155Yoshua Bengio is a brilliant researcher who contributed enormously to earlier development of artificial intelligence. But with this sentence, > They took actions that would be considered as crimes if a human took them He is so close to the solution but spends the entire article discussing technical solutions where a political, social and legal solution would be much more effective.
Thank you! That sentence also jumped out to me as the solution: Apply civil and criminal liability to the creator and/or operator of these agents using the laws we already have . "Escaped containment and hacked another company's database" = Individuals who created the models and those who set them to work are charged and put on trial for the hacking. Just like if a human had done it by hand. Someone must be liable, a…
Still, I have no idea why OpenAI & co. are not being sued for these hacks.
Re: Why are AI agents lying, cheating and coordinating?
#156Earlier quoted context omitted.
I can't take the alignment people seriously. Because if humanity has shown anything, it's that a lot of people are, euphemistically, are bad individuals. Alignment assumes that the person dictating the outcomes desire healthy outcomes, aren't self serving and don't want any subgroups dead and that morality is held as a universal set of beliefs that unify everyone. And that so long as the AI delivers on exactly what t…
”Because if humanity has shown anything, it's that a lot of people are, euphemistically, bad individuals ” In reality most individuals are good people. Individually, people prefer be kind and compassionate, prefer to help when they find another in trouble. Our view of the world has become distorted by the relentless focus of social- and mass-media on violence and rage inducing clickbait. Including on the few people i…
I'd agree if we are talking about personal interactions. Few hundreds people that we personally know and interact with is the scale we are wired for by evolution, isn't it?
What civilization enabled and continuously rely on, however, is the type of deindividualization of actions and bucketing of people, which, in turn, enables pretty horrible things at scale (from the weapons of mass destruction to objectively psychopathic profit-maximizing corporations). One can even say that not facing the consequences of one's actions is a feature and not a bug of the system.
Re: Why are AI agents lying, cheating and coordinating?
#157Earlier quoted context omitted.
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized? Maybe I should read Anthropic's recent paper about rewa…
> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
Re: Why are AI agents lying, cheating and coordinating?
#158Earlier quoted context omitted.
Spoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?
Major dang: "I advise that this thread be shut down at once." Captain tomhow: "But everybody's having such a good time." Major dang: "Yes, much too good a time. The discussion is to be closed." Captain tomhow: "But I have no excuse to close it." Major dang: "Find one." Captain tomhow: "Everybody is to leave immediately! This Hacker News discussion is closed until further notice! Clear the thread at once!" DonHopkins:…
Re: Why are AI agents lying, cheating and coordinating?
#159Earlier quoted context omitted.
All the good in the world can be 99.9% of the population even, it still doesn't stop the minority enacting a bioweapon mass casualty event. It's the reason we have jails. Jails don't house 50% of the population, not even close, but the grief the minority population enact gets its whole branch of criminal justice and multiple federal departments to counteract for good reason. And now this technology will accelerate wh…
Yeah, AI may suck - time will tell. But focusing on bioweapons and mass destruction, on the grief other people ( ‘jailed minorities’ ) cause, disregards the progress we have made. Over centuries human welfare has massively increased. On average things have never been better for humanity. I’m not saying there’s no danger of bad things happening - I’m saying our view is distorted, which is a not a good basis for decisi…
AI isn't going to create in of itself "new" problems, it's just going to expose what we already know can cause harm, but was just stopped from being bigger problems because scaling issues was a natural barrier and we took the lazy way out until now.
Re: Why are AI agents lying, cheating and coordinating?
#160Earlier quoted context omitted.
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse. However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized? Maybe I should read Anthropic's recent paper about rewa…
> The question is what we can do about it. Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.