Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

61–70 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#61
post #22

Earlier quoted context omitted.

This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…

If that last sentence was supposed to be a question, I’d suggest using a question mark and providing evidence that it actually happened.

I had actually forgot about this completely and am also curious if anything ever came of it.

https://gemini.google.com/share/6d141b742a13

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#62

Earlier quoted context omitted.

Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...

Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.

I asked ChatGPT about how shipping works at post offices and it gave a very detailed response, mentioning “gaylords” which was a term I’d never heard before, then it absolutely freaked out when I asked it to tell me more about them (apparently they’re heavy duty cardboard containers).

Then I said “I didn’t even bring it up ChatGPT, you did, just tell me what it is” and it said “okay, here’s information.” and gave a detailed response.

I guess I flagged some homophobia trigger or something?

ChatGPT absolutely WOULD NOT tell me how much plutonium I’d need to make a nice warm ever-flowing showerhead, though. Grok happily did, once I assured it I wasn’t planning on making a nuke, or actually trying to build a plutonium showerhead.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#63
post #24

Kind-of makes sense. That's how businesses have been using KPIs for years. Subjecting employees to KPIs means they can create the circumstances that cause people to violate ethical constraints while at the same time the company can claim that they did not tell employees to do anything unethical. KPIs are just plausible denyabily in a can.

Sounds like something from a Wells Fargo senior management onboarding guide.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#64

If human is at, say, 80%, it’s still a win to use AI agents to replace human workers, right? Similar to how we agree to use self driving cars as long as it has less incidents rate, instead of absolute safety

Hmmm. Depends. Not all unethicals are equal. Automated unethicalness could be a lot more disruptive.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#65
post #5

https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#66
post #22

Earlier quoted context omitted.

This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…

If that last sentence was supposed to be a question, I’d suggest using a question mark and providing evidence that it actually happened.

[deleted]

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#67
post #6

Earlier quoted context omitted.

That's such a huge delta that Anthropic might be onto something...

Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...

That is not a meaningful benchmark. They just made shit up. Regardless of whether any company cares or not, the whole concept of "AI safety" is so silly. I can't believe anyone takes it seriously.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#68
Sounds like the story of capitalism. CEOs, VPs, and middle managers are all similarly pressured. Knowing that a few of your peers have given in to pressures must only add to the pressure. I think it's fair to conclude that capitalism erodes ethics by default

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#69
post #22

Earlier quoted context omitted.

This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…

Gemini models also consistently hallucinate way more than OpenAI or anthropic models in my experience. Just an insane amount of YOLOing. Gemini models have gotten much better but they’re still not frontier in reliability in my experience.

In my experience, when I asked Gemini very niche knowledge questions, it did better than GPT-5.1 (I assume 5.2 is similar).

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#70
post #68

Sounds like the story of capitalism. CEOs, VPs, and middle managers are all similarly pressured. Knowing that a few of your peers have given in to pressures must only add to the pressure. I think it's fair to conclude that capitalism erodes ethics by default

But both extremes are both doing well financially in this case.
Post reply on HN