Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

301–310 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#301

Earlier quoted context omitted.

Honestly for research level math, the reasoning level of Gemini 3 is much below GPT 5.2 in my experience--but most of the failure I think is accounted for by Gemini pretending to solve problems it in fact failed to solve, vs GPT 5.2 gracefully saying it failed to prove it in general.

Have you tried Deep Think? You only get access with the Ultra tier or better... but wow. It's MUCH smarter than GPT 5.2 even on xhigh. It's math skills are a bit scary actually. Although it does tend to think for 20-40 minutes.

I tried Gemini 2.5 Deep Think, was not very impressed ... too much hallucinations. In comparison GPT 5.2 extended time hallucinates at like <25% of the time and if you ask another copy to proofread it goes even lower.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#302

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

I think this also shows up outside an AI safety or ethics framing and in product development and operations. Ultimately "judgement," however you wish to quantify that fuzzy concept, is not purely an optimization exercise. It's far more a probabilistic information function from incomplete or conflicting data. In product management (my domain), decisions are made under conflicting constraints: a big customer or account…

As Goodhart's law states, "When a measure becomes a target, it ceases to be a good measure". From an organizational management perspective, one way to partially work around that problem is by simply adding more measures thus making it harder for a bad actor to game the system. The Balanced Scorecard system is one approach to that.

https://balancedscorecard.org/

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#303
post #294
post #234

Earlier quoted context omitted.

What you have quoted says a third of people who thought it was real didn’t disobey the experimenter when they thought they were delivering dangerous and lethal electric shocks to a human. Is that correct?

Maybe there was an edit but it's the opposite, 66% disobeyed.

Right, so a third didn’t disobey.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#304
post #302

Earlier quoted context omitted.

I think this also shows up outside an AI safety or ethics framing and in product development and operations. Ultimately "judgement," however you wish to quantify that fuzzy concept, is not purely an optimization exercise. It's far more a probabilistic information function from incomplete or conflicting data. In product management (my domain), decisions are made under conflicting constraints: a big customer or account…

As Goodhart's law states, "When a measure becomes a target, it ceases to be a good measure". From an organizational management perspective, one way to partially work around that problem is by simply adding more measures thus making it harder for a bad actor to game the system. The Balanced Scorecard system is one approach to that. https://balancedscorecard.org/

Agreed, Goodhart’s Law captures the failure mode well intentioned KPIs and OKRs may miss, let alone agentic automation

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#305
post #302

Earlier quoted context omitted.

I think this also shows up outside an AI safety or ethics framing and in product development and operations. Ultimately "judgement," however you wish to quantify that fuzzy concept, is not purely an optimization exercise. It's far more a probabilistic information function from incomplete or conflicting data. In product management (my domain), decisions are made under conflicting constraints: a big customer or account…

As Goodhart's law states, "When a measure becomes a target, it ceases to be a good measure". From an organizational management perspective, one way to partially work around that problem is by simply adding more measures thus making it harder for a bad actor to game the system. The Balanced Scorecard system is one approach to that. https://balancedscorecard.org/

This extends beyond AI agents. I'm seeing it in real time at work — we're rolling out AI tools across a biofuel brokerage and the first thing people ask is "what KPIs should we optimize with this?"

The uncomfortable answer is that the most valuable use cases resist single-metric optimization. The best results come from people who use AI as a thinking partner with judgment, not as an execution engine pointed at a number.

Goodhart's Law + AI agents is basically automating the failure mode at machine speed.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#306

Earlier quoted context omitted.

How is a corporation "immortal"? What is the oldest corporation in the world? I mean, aside from churches and stuff. Corporations can die or be killed in numerous ways. Not many of them will live forever. Most will barely outlive a normal human's lifespan. By definition, since a corporation comprises a group of people, it could never outlive the members, should they all die at some point. Let us also draw a distincti…

What needs to do a company from fortune 7 to die? If kills 1 person they won’t close Google. If steals 1 billion, won’t close either. So what needs to do such a company to be closed down? I think it’s almost impossible to shut down

Look to history. Here's a list of "Fortune 7" companies from about 50 years ago.

IBM

AT&T

Exxon

General Motors

General Electric

Eastman Kodak

Sears, Roebuck & Co.

Some of them died. Others are still around but no longer in the top 7. Why is that? Eventually every high-growth company misses a disruptive innovation or makes a key strategic error.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#307

Earlier quoted context omitted.

The fact that the guy leading the development of Gemini was on Epstein's island is probably unrelated.

I can't find anything verifiable related to your statement ...

https://en.wikipedia.org/wiki/Prominent_individuals_mentione...

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#308

Earlier quoted context omitted.

Where is Pratchett when we need him? I wonder how he would have chose to anthropomorphize anthropomorphism. A sort of meta anthropomorphization.

I’m certainly no Pratchett, so I can’t speak to that. I would say there’s an enormous round coin upon which sits an enormous giant holding a magnifying glass, looking through it down at her hand. When you get closer, you see the giant is made of smaller people gazing back up at the giant through telescopes. Get even closer and you see it’s people all the way down. The question of what supports the coin, I’ll leave to…

> We anthropomorphize because we’re people, and it’s people all the way down.

Nice bit of writing. Wish I had more than one upvote to give.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#309

Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked. This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and s…

What kind of value do you get from talking to it about “sensitive” subjects? Speaking as someone who doesn’t use AI, so I don’t really understand what kind of conversation you’re talking about

One example - I'm doing research for some fiction set in the late 19th century, when strychnine was occasionally used as a stimulant. I want to understand how / when it would have been used and dosages, and ChatGTP shut down that conversation "for safety".
Post reply on HN