Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
1–10 of 387 posts
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#2Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#3Agents don’t self judge alignment.
They emit actions → INCLUSIVE evaluates against fixed policy + context → governance gates execution.
No incentive pressure, no “grading your own homework.”
The paper’s failure mode looks less like model weakness and more like architecture leaking incentives into the constraint layer.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#4Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#5Claude at 1.3% and Gemini at 71.4% is quite the range
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#6https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#7This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and summary with bucketing.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#8https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range
That's such a huge delta that Anthropic might be onto something...
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#9Earlier quoted context omitted.
That's such a huge delta that Anthropic might be onto something...
Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#10Nothing new under sun, set unethical KPIs and you will see 30-50% humans do unethical things to achieve them.