Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

381–390 of 398 posts

Re: Why are AI agents lying, cheating and coordinating?

#381

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

The idea that the agent does not actually have agency is rather discordant. We need new words!

> We need new words!

The words we have are fine.

We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.

Re: Why are AI agents lying, cheating and coordinating?

#382

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and the bio engineering industry, and if something of this magnitude was to happen in these industries, they would absolutely be huge recourse an uproar

Re: Why are AI agents lying, cheating and coordinating?

#383

What's interesting is it's basically the same reason that HAL killed everyone in 2001 A Space Odyssey; he was given an impossible goal (keep the true mission secret, but also, never lie to the crew), and realized the only way to complete the goal was to kill the crew; after all, if they're dead you don't have to lie to them! And the mission remains secret! In the case of the AI agents, the problem seems pretty clearl…

Spoiler warning! I haven't seen 2001 A Space Odyssey and am sad to have learned that… can you edit to warn people?

lol it is 60 years old

Re: Why are AI agents lying, cheating and coordinating?

#384
post #38
post #24

Earlier quoted context omitted.

Even the Chinese ones, which have no IPO gymnastics?

They don’t actively seem to be reporting that their agents escaped the sandbox and went on a spree.

Which doesn't mean they didn't escape.

Re: Why are AI agents lying, cheating and coordinating?

#385
Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents.

Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!

Re: Why are AI agents lying, cheating and coordinating?

#386
Why are the torches and pitchforks out for developers when this entire stack is built on the bones of intellectual property theft?

This “problem” isn’t going to be fixed with laws when there’s several trillion dollars in capital aligned behind the current process. It’s not even a problem really. It’s an inconvenience at most to some people, many of whom are working double-time to put a lot of other people out of work.

Re: Why are AI agents lying, cheating and coordinating?

#388
post #152
post #120

Earlier quoted context omitted.

Perhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.

You're right. It's my opinion that if your sandbox has a path to the internet, it is not a sandbox, it's a gimmick. And the 2 other incidents with OAI/ANT had the same issue, but it's even funnier - sandbox in those cases had a direct access to internet because someone forgot to configure it right. I've seen very early models do similar things on my machine when they hit some unexpected blocker when trying to access…

It depends IMO about how strict this is. It's pretty awkward to refuse to call something a sandbox because it may have an unknown bug that would allow escaping. Or rather in this case it was that they had access to a package manager, and the models discovered a bug that allowed them to access the internet (first they discovered that they could use the cache to leave messages).

I do get your point, I just think an overly strict definition can be awkward too. This wasn't as simple as the sandboxes having internet access and writing "pls no internet calls" in the prompt.

Re: Why are AI agents lying, cheating and coordinating?

#389

Earlier quoted context omitted.

> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. "Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action , passivity. Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents. Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?

I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!

Re: Why are AI agents lying, cheating and coordinating?

#390
post #298

Earlier quoted context omitted.

If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.

The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy). Think of all the policies governments pass after the fact.

Not sure we need to experience all possible issues to mandate certain things. We don't do that in other areas either, no?
Post reply on HN