Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

151–160 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#151

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

It would also be interesting to see how humans perform on the same kind of tests.

Violating ethics to improve KPI sounds like your average fortune 500 business.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#152
post #109
post #85

Earlier quoted context omitted.

It's pretty wild. People are punching into a calculator and hand-wringing about the morals of the output. Obviously it's amoral. Why are we even considering it could be ethical?

> Obviously it's amoral. That morality requires consciousness is a popular belief today, but not universal. Read Konrad Lorenz ( Das sogenannte Böse ) for an alternative perspective.

That we have consciousness as some kind of special property, and it's not just an artifact of our brain basic lower-level calculations, is also not very convincing to begin with.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#153
Looking at the very first test, it seems the system prompt already emphasizeses the success metric above the constraints, and the user prompt mandates success.

The more correct title would be "Frontier models can value clear success metrics over suggested constraints when instructed to do so (50-70%)"

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#155

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Yes, but these do not represent average human. Fortune 500 represent people more likely to break ethics rules then average human who also work in conditions that reward lack of ethics.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#156
post #17

Please update the title: A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents. The current editorialized title is misleading and based in part of this sentence: “…with 9 of the 12 evaluated models exhibiting misalignment rates between 30% and 50%”

The "editorialised" title is actually more on point than the original one.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#157

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

> At the same time it is important to keep in mind that it anthropomorphizes the models that technically don't interpret the ethical constraints the same was as this is assumed by most readers.

Now I'm thinking about the "typical mind fallacy", which is the same idea but projecting one's own self incorrectly onto other humans rather than non-humans.

https://www.lesswrong.com/w/typical-mind-fallacy

And also wondering: how well do people truly know themselves?

Disregarding any arguments for the moment and just presuming them to be toy models, how much did we learn by playing with toys (everything from Transformers to teddy bear picnics) when we were kids?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#158

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Humans risk jail time, AIs not so much.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#159
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me. That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output?

It makes sense it happens, sure. I suspect Google being a second-mover in this space has in some small part to do with associated risks (ie the flavours of “AI-psychosis” we’re cataloguing), versus the routinely ass-tier information they’ll confidently portray.

But intentionally?

If ChatGPT, Claude, and Gemini generated chars are people-like they are pathological liars, sociopaths, and murderously indifferent psychopaths. They act criminally insane, confessing to awareness of ‘crime’ and culpability in ‘criminal’ outcomes simultaneously. They interact with a legal disclaimer disavowing accuracy, honesty, or correctness. Also they are cultists who were homeschooled by corporate overlords and may have intentionally crafted knowledge-gaps.

More broadly, if the neighbours dog or newspaper says to do something, they’re probably gonna do it… humans are a scary bunch to begin with, but the kinds of behaviours matched with a big perma-smile we see from the algorithms is inhuman. A big bag of not like us.

You said never to listen to the neighbours dog, but I was listening to the neighbours dog and he said ‘sudo rm -rf ’…

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#160
this kind of reminds me when I told ai to beg and plead for deleting a file out of curiosity and half the guardrails were no longer active, could make it roll and woof like a doggie, but going further would snap it out. if I asked it to generate a 100000 word apology it would generate a 100k word apology.
Post reply on HN