Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

211–220 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#211
post #208

Earlier quoted context omitted.

Not quite. The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). However, the vast majority of psychological research over the last 80 years heavily favours a situational explanation (it's about the environment/system). Everyone (in the field) got really in…

> The Milgram and Stanford Prison experiments are the most obvious examples. BOTH are now considered bad science. BOTH are now used as examples of "how not to do the science". > The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). I did not said nor implie…

Not sure where you get that for Milgram. That's been replicated lots of times, in different countries, with different compositions of people, and found to be broadly replicable. Burger in '09, Sheridan & King in '72, Dolinski and co in '17, Caspar in '16, Haslam & Reicher which I referenced somewhere else in the thread...

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#212
post #183

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

A remarkable number of humans given really quite basic feedback will perform actions they know will very directly hurt or kill people. There are a lot of critiques about quite how to interpret the results but in this context it’s pretty clear lots of humans can be at least coerced into doing something extremely unethical. Start removing the harm one, two, three degrees and add personal incentives and is it that surpr…

> 2012, Australian psychologist Gina Perry investigated Milgram's data and writings and concluded that Milgram had manipulated the results, and that there was a "troubling mismatch between (published) descriptions of the experiment and evidence of what actually transpired." She wrote that "only half of the people who undertook the experiment fully believed it was real and of those, 66% disobeyed the experimenter".[29][30] She described her findings as "an unexpected outcome" that

Its unlikely Milligram played am unbiased role in, if not the sirext cause of the results.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#214

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

Quite possibly, workable ethics will pretty much require full-fledged General Artificial Intelligence, verging on actual Self-Awareness.

There's a great discussion of this in the (Furry) web-comic Freefall:

http://freefall.purrsia.com/

(which is most easily read using the speed reader: https://tangent128.name/depot/toys/freefall/freefall-flytabl... )

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#215

Earlier quoted context omitted.

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Humans risk jail time, AIs not so much.

From an IBM training manual (1979):

>A computer can never be held accountable

>Therefore a computer must never make a management decision

The (EDITED) corollary would arguably be:

>Corporations are amoral entities which are potentially immortal who cannot be placed behind bars. Therefore they should never be given the rights of human beings.

(potentially, not absolutely immortal --- would wording as "not mortal by essence/nature"? be better?)

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#216

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

From an IBM training manual (1979): >A computer can never be held accountable >Therefore a computer must never make a management decision The (EDITED) corollary would arguably be: >Corporations are amoral entities which are potentially immortal who cannot be placed behind bars. Therefore they should never be given the rights of human beings. (potentially, not absolutely immortal --- would wording as "not mortal by es…

How is a corporation "immortal"?

What is the oldest corporation in the world? I mean, aside from churches and stuff.

Corporations can die or be killed in numerous ways. Not many of them will live forever. Most will barely outlive a normal human's lifespan.

By definition, since a corporation comprises a group of people, it could never outlive the members, should they all die at some point.

Let us also draw a distinction between the "human being" and the "person". A corporation is granted "personhood" but this is not equivalent to "humanity". Being composed of humans, the members of any corporation collectively enjoy their individual rights in most ways.

A "corporate person" is distinct from a "human person", and so we can recognize that "corporate rights" are in a different category, and regulate accordingly.

A corporation cannot be "jailed" but it can be fined, it can be dissolved, it can be sanctioned in many ways. I would say that doing business is a privilege and not a right of a corporation. It is conceivable that their ability to conduct business could be restricted in many ways, such as local only, or non-interstate, or within their home nation. I suppose such restrictions could be roughly analogous to being "jailed"?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#218

Earlier quoted context omitted.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me. That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output? It makes sense it happens, sure. I suspect Google being a second-mover in this space has in some small part to do with associated risks (ie the flavours of “AI-psychosis” we’re cataloguing), versus the routinely ass-tier information they’ll confidently portray. But intentionally? If ChatGPT, Claud…

>It makes sense to attribute human characteristics or behaviour to a non-reasoning data-set-constrained algorithms output?

It makes total sense, since the whole development of those algorithms was done so that we get human characteristics and behaviour from them.

Not to mention, your argument is circular, amounting to that an algorithm can't have "human characteristics or behaviour" because it's an algorithm. Describing them as "non reasoning" is already begging the question, as any any naive "text processing can't produce intelligent behavior" argument, which is as stupid as saying "binary calculations on 0 and 1 can't ever produce music".

Who said human mental processing itself doesn't follow algorithmic calculations, that, whatever the physical elements they run on, can be modelled via an algorithm? And who said that algorithm won't look like an LLM on steroids?

That the LLM is "just" fed text, doesn't mean it can get a lot of the way to human-like behavior and reasoning already (being able to pass the canonical test for AI until now, the Turing test, and hold arbitrary open ended conversations, says it does get there).

>If ChatGPT, Claude, and Gemini generated chars are people-like they are pathological liars, sociopaths, and murderously indifferent psychopaths. They act criminally insane, confessing to awareness of ‘crime’ and culpability in ‘criminal’ outcomes simultaneously. They interact with a legal disclaimer disavowing accuracy, honesty, or correctness. Also they are cultists who were homeschooled by corporate overlords and may have intentionally crafted knowledge-gaps.

Nothing you wrote above doesn't apply to more or less the same degree to humans.

You think humans don't do all mistakes and lies and hallucination-like behavior (just check the bibliography on the reliability of human witnesses and memory recall)?

>More broadly, if the neighbours dog or newspaper says to do something, they’re probably gonna do it… humans are a scary bunch to begin with, but the kinds of behaviours matched with a big perma-smile we see from the algorithms is inhuman. A big bag of not like us.

Wishful thinking. Tens of millions of AIs didn't vote Hitler to power and carried the Holocaust and mass murder around Europe. It was German humans.

Tens of millions of AIs didn't have plantation slavery and seggregation. It was humans again.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#219
I'm noticing an increasing desire in some businesses for plausibly deniable sociopathy. We saw this with the Lean Startup movement and we may see an increasing amount in dev shops that lean more into LLMs.

Trading floors are an established example of this, where the business sets up an environment that encourages its staff to break the rules while maintaining plausible deniability. Gary's economics references this in an interview where he claimed Citigroup were attempting to threaten him with all the unethical things he'd done with such confidence that he had, only to discover he hadn't.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#220
post #85
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

It's pretty wild. People are punching into a calculator and hand-wringing about the morals of the output. Obviously it's amoral. Why are we even considering it could be ethical?

Have you tried "kill all the poor?" [0]

[0] https://www.youtube.com/watch?v=s_4J4uor3JE

Post reply on HN