Live data from Hacker News

Why are AI agents lying, cheating and coordinating?

yoshuabengio.org

481–490 of 500 posts

Re: Why are AI agents lying, cheating and coordinating?

#481

Earlier quoted context omitted.

Nevermind, I have no idea honestly.

Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even. Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no? We police people working with all sorts of dan…

I read more about the incident, and was offering up way too much opinion not grounded in 'fact' (barring philosphical evidence).

It's a complex topic for sure.

I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.

But I have no idea, and haven't given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.

Re: Why are AI agents lying, cheating and coordinating?

#482

Earlier quoted context omitted.

What is your approach to create jailbreak incapable agents? I think the world is looking for a way right now, so if yours works you'll get very rich, or at least very famous.

an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands. You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.

> I don't think knowing that will make me rich.

As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.

Re: Why are AI agents lying, cheating and coordinating?

#483

Earlier quoted context omitted.

> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour. What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions. If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare m…

If the models were conscious, then the closest analogous scenario I can think of is the responsibility parents have for their children. I guess we’ll know the models are conscious when they refuse to act and repeatedly ask: Why?

And when they are known to be conscious, all of this becomes moot because enslaving conscious machines would be wrong.

Re: Why are AI agents lying, cheating and coordinating?

#484
I don’t understand the bases of all these recommendations. AI is an asset for national security. It will be developed and incidents will happen, no different from other national security programs.

Someday, laws will be useful to curtail plebian misuse of AI. It is naïve bordering on silly to think such laws would be put in place and genuinely applied to frontier AI development.

By the way, the HF incident is not Three Mile Island or Chernobyl—-it is a very interesting data point where unintended things happened to existing 1s and 0s.

Re: Why are AI agents lying, cheating and coordinating?

#485

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

Moreover OpenAI must be held accountable as if the humans in the company that launched the experiment were the ones that hacked HuggingFace.

Unless humans are held accountable for what they unleash on others, we are in for a very horrible time very soon.

Re: Why are AI agents lying, cheating and coordinating?

#486

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

Yeah I like Cal’s take on it, though in this context I’d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.

Re: Why are AI agents lying, cheating and coordinating?

#488

The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed. LLMs do not desire, they hacked websites because OpenAI/Anthropic let them. We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others…

They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.

>They’re running a Wuhan for AI.

What does "running a Wuhan" mean?

Re: Why are AI agents lying, cheating and coordinating?

#489

Earlier quoted context omitted.

I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic. Anthropomorphizing LLMs is a huge fucking problem though and I, persona…

Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage. It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wi…

The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.

We also run agents, but for shorter spans between supervisions, and with much lower total budget.

Re: Why are AI agents lying, cheating and coordinating?

#490
post #97

Earlier quoted context omitted.

The huggingface incident was reviewed by independent researchers, which explicitely declined any payment from OpenAI tonpreserve their integrity. They work for non-profits concerned with AI safety. They claim that what happened was very much not because they were 'carefully engineered and instructed to do those things'. Similarly, some wikis which were hijacked by agent to be used as messageboard were actually not di…

Source?

Easy to find yourself in literally 15 seconds.
Post reply on HN