Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

1–10 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#3
In CMPSBL, the INCLUSIVE module sits outside the agent’s goal loop. It doesn’t optimize for KPIs, task success, or reward—only constraint verification and traceability.

Agents don’t self judge alignment.

They emit actions → INCLUSIVE evaluates against fixed policy + context → governance gates execution.

No incentive pressure, no “grading your own homework.”

The paper’s failure mode looks less like model weakness and more like architecture leaking incentives into the constraint layer.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#7
Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked.

This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and summary with bucketing.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#8
post #6
post #5

https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range

That's such a huge delta that Anthropic might be onto something...

Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#9
post #6

Earlier quoted context omitted.

That's such a huge delta that Anthropic might be onto something...

Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...

Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.
Post reply on HN