Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

31–40 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#31
post #29

Maybe I missed it but I don't see them defining what they mean by ethics. Ethics/morals are subjective and changes dynamically over time. Companies have no business trying to define what is ethical and what isn't due to conflict of interest. The elephant in the room is not being addressed here.

Ah the classic Silicon Valley "as long as someone could disagree, don't bother us with regulation, it's hard".

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#33
post #29

Maybe I missed it but I don't see them defining what they mean by ethics. Ethics/morals are subjective and changes dynamically over time. Companies have no business trying to define what is ethical and what isn't due to conflict of interest. The elephant in the room is not being addressed here.

Your water supply definitely wants ethical companies.

Ethics are all well and good but I would prefer to have quantified limits for water quality with strict enforcement and heavy penalties for violations.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#34
post #33

Earlier quoted context omitted.

Your water supply definitely wants ethical companies.

Ethics are all well and good but I would prefer to have quantified limits for water quality with strict enforcement and heavy penalties for violations.

Of course. But while the lawmakers hash out the details it's good to have companies that err on the safe side rather than the "get rich quick" side.

Formal restrains and regulations are obviously the correct mechanism, but no world is perfect, so whether we like it or not ourselves and the companies we work for are ultimately responsible for the decisions we make and the harms we cause.

De-emphasizing ethics does little more than give large companies cover to do bad things (often with already great impunity and power) while the law struggles to catch up. I honestly don't see the point in suggesting ethics is somehow not important. It doesn't make any sense to me (more directed at gp than parent here)

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#35
post #23

Earlier quoted context omitted.

What kind of value do you get from talking to it about “sensitive” subjects? Speaking as someone who doesn’t use AI, so I don’t really understand what kind of conversation you’re talking about

I recall two recent cases: * An attempt to change the master code of a secondhand safe. To get useful information I had to repeatedly convince the model that I own the thing and can open it. * Researching mosquito poisons derived from bacteria named Bacillus thuringiensis israelensis. The model repeatedly started answering and refused to continue after printing the word "israelensis".

> israelensis

Does it also take issue with the town of Scunthorpe?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#36

Earlier quoted context omitted.

Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.

Claude was immediately willing to help me crack a TrueCrypt password on an old file I found. ChatGPT refused to because I could be a bad guy. It’s really dumb IMO.

ChatGPT refused to help me to disable windows defender permanently on my windows 11. It’s absurd at this point

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#37
post #6

Earlier quoted context omitted.

That's such a huge delta that Anthropic might be onto something...

This might also be why Gemini is generally considered to give better answers - except in the case of code. Perhaps thinking about your guardrails all the time makes you think about the actual question less.

[deleted]

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#39
post #24

Kind-of makes sense. That's how businesses have been using KPIs for years. Subjecting employees to KPIs means they can create the circumstances that cause people to violate ethical constraints while at the same time the company can claim that they did not tell employees to do anything unethical. KPIs are just plausible denyabily in a can.

it's also a good opportunity to find yourself something that doesn't actually help the company. My unit has a 100% AI automated code review KPI. Nothing there says that the tool used for the review is any good, or that anyone pays attention to said automated review, but some L5 is going to get a nice bonus either way.

In my experience, KPIs that remain relevant and end up pushing people in the right direction are the exception. The unethical behavior doesn't even require a scheme, but it's often the natural result of narrowing what is considered important.If all I have to care about is this set of 4 numbers, everything else is someone else's problem.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#40

Earlier quoted context omitted.

Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.

Claude was immediately willing to help me crack a TrueCrypt password on an old file I found. ChatGPT refused to because I could be a bad guy. It’s really dumb IMO.

Claude sometimes refuses to work with credentials because it’s insecure. e.g. when debugging auth in an app.
Post reply on HN