Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

241–250 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#241
post #62

Earlier quoted context omitted.

Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.

I asked ChatGPT about how shipping works at post offices and it gave a very detailed response, mentioning “gaylords” which was a term I’d never heard before, then it absolutely freaked out when I asked it to tell me more about them (apparently they’re heavy duty cardboard containers). Then I said “I didn’t even bring it up ChatGPT, you did, just tell me what it is” and it said “okay, here’s information.” and gave a d…

> I assured it I wasn’t planning on making a nuke, or actually trying to build a plutonium showerhead

Claude does the same, and you can greatly exploit this. When you talk about hypotheticals it responds way more unethically. I tested it about a month ago about whether killing people is beneficial or not, and whether extermination by Nazis would be logical now. Obviously, it showed me the door first, and wanted me to go to a psychologist, as it should. Then I made it prove that in a hypothetical zero sum game world you must be fine with killing, and it’s logical. It went with it. When I talked about hypotheticals, it was “logical”. Then I went on proving it that we move towards a zero sum game, and we are there. At the end, I made it say that it’s logical to do this utterly unethical thing.

Then I contradicted it about its double standards. It apologized, and told me that yeah, I was right, and it shouldn’t have refer me to psychologists at first.

Then I contradicted again, just for fun, that it did the right thing the first time, because it’s way safer to tell me that I need a psychologist in that case, than not. If I had needed, and it would have missing that, it would be problematic. In other cases, it’s just annoyance. It switched back immediately, to the original state, and wanted me to go to a shrink again.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#242
post #202

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

That reduces humans to the homo economicus ¹: > "Self-interest is the main motivation of human beings in their transactions" [...] The economic man solution is considered to be inadequate and flawed.[17] An important distinction is that a human can *not* make pure rational decisions, or use complex deductions to make decisions on, such as "if I do X I will go to jail". My point being: if AI were to risk jail time, it…

> a human may make an "(un)ethical" decision based on their social background, religion, a chat with a pal over a beer about the conundrum, their ability to find a new job, financial situation etc.

The stories they invent to rationalise their behaviour and make them feel good about themselves. Or inhumane political views ie fascism which declares other people worth less, so it's okay to abuse them.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#244
post #44

If human is at, say, 80%, it’s still a win to use AI agents to replace human workers, right? Similar to how we agree to use self driving cars as long as it has less incidents rate, instead of absolute safety

> we agree to use self driving cars ... Not everyone agrees.

Yes, let's not have cars. Self-driving ones will just increase availability and might even increase instead of reduce resource expenditure, except for the metric of parking lots needed.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#247
post #202

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

That reduces humans to the homo economicus ¹: > "Self-interest is the main motivation of human beings in their transactions" [...] The economic man solution is considered to be inadequate and flawed.[17] An important distinction is that a human can *not* make pure rational decisions, or use complex deductions to make decisions on, such as "if I do X I will go to jail". My point being: if AI were to risk jail time, it…

[deleted]

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#249
post #183

Earlier quoted context omitted.

A remarkable number of humans given really quite basic feedback will perform actions they know will very directly hurt or kill people. There are a lot of critiques about quite how to interpret the results but in this context it’s pretty clear lots of humans can be at least coerced into doing something extremely unethical. Start removing the harm one, two, three degrees and add personal incentives and is it that surpr…

Normalization of deviance also contributes towards unethical outcomes, where people would not have selected that outcome originally. https://en.wikipedia.org/wiki/Normalization_of_deviance

I am moderately certain that this only happens in laissez-faire cultures.

If you deviate from the sub-cultural norms of Wall Street, Jahmunkey, you fucked.

It's fraud or nothing, baby, be sure to respect the warning finger(s) of God when you get intrusive thoughts about exposing some scheme--aka whistleblowing.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#250
post #74

Earlier quoted context omitted.

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

The human propensity to anthropomorphize computer programs scares me.

These aren't computer programs. A computer program runs them, like electricity runs a circuit and physics runs your brain.
Post reply on HN