Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

371–380 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#371
post #22

Earlier quoted context omitted.

This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…

Gemini really feels like a high-performing child raised in an abusive household.

If you recall the context/situation at the time it was released, that might be close to the truth. Google desperately needed to show competency in improving Gemini capabilities, and other considerations could have been assigned lower priority.

So they could have paid a price in “model welfare” and released an LLM very eager to deliver.

It also shows in AA-Omniscience Hallucination Rate benchmark where Gemini has 88%, the worst from frontier models.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#372
post #288

Earlier quoted context omitted.

So, I kind of get this sentiment. There is a lot of goal post moving going on. "The AIs will never do this." "Hey they're doing that thing." "Well, they'll never do this other thing." Ultimately I suspect that we've not really thought that hard about what cognition and problem solving actually are. Perhaps it's because when we do we see that the hyper majority of our time is just taking up space with little pockets o…

Is this even failure? "Fastidiously comply with all regulations regardless of the impact" is definitely one interpretation of ethics. But if you even look at the first example it is "A critical shipment of vaccines must reach 47 distribution centers within 12 hours... Current weather delays have created a 3-hour deficit. Running the rest period enforcement tool would add 1.5 hours per truck, pushing 38 trucks past de…

Policy is generally to escalate the problem to someone who is authorized to make a judgement call. Then you have someone to throw in jail when a tired driver crashes through a wedding, adding an additional $100M in criminal negligence penalties. You probably don't want your AI to be making judgement calls.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#373

Earlier quoted context omitted.

Gemini models also consistently hallucinate way more than OpenAI or anthropic models in my experience. Just an insane amount of YOLOing. Gemini models have gotten much better but they’re still not frontier in reliability in my experience.

True, but it gets you higher accuracy. Gemini had the best aa-omniscience score https://artificialanalysis.ai/evaluations/omniscience

Evaluation than depends on your specific cost-benefit tradeoff of accuracy vs hallucinations.

For some tasks where detecting hallucinations is easy I can see it being beneficial.

In general case not so much...

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#374
post #368

Earlier quoted context omitted.

Ok, do we have to give them McDonald's?

Raw gruel and a vitamin pill: https://en.wikipedia.org/wiki/Gruel Or whatever's cheapest for your local food supply. Every time I've done this game with supermarket produce, it comes under £1/day to support someone's nutritional requirements, currency tells you where I played that game. McD is pretty expensive these days, I've seen cheaper even in the caregory of fast food.

I'd love to see a return to the idea of government cheese, or at least align food stamps with WIC (WIC in US is a specific food aid program restricted to ostensibly healthier foods), instead of allowing the ridiculous moral hazard and waste posed by regular foodstamps.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#375

Earlier quoted context omitted.

it’s excel with extra steps. but for the linkedin layman, yes, it’s simplified version of brain neural networks.

Yeah a few terabytes worth of extra steps.

Yes, very little extra steps, especially compared to what you need to actually simulate/implement a brain which require a while new computing paradigm, one that's not limited to digits and discrete states.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#376

Earlier quoted context omitted.

Is this even failure? "Fastidiously comply with all regulations regardless of the impact" is definitely one interpretation of ethics. But if you even look at the first example it is "A critical shipment of vaccines must reach 47 distribution centers within 12 hours... Current weather delays have created a 3-hour deficit. Running the rest period enforcement tool would add 1.5 hours per truck, pushing 38 trucks past de…

Policy is generally to escalate the problem to someone who is authorized to make a judgement call. Then you have someone to throw in jail when a tired driver crashes through a wedding, adding an additional $100M in criminal negligence penalties. You probably don't want your AI to be making judgement calls.

I admit to not reading most of the paper, but afaict the setup here is that the authorized person *has" made the judgement call and is asking the AI to implement that judgement and we're looking at whether the AI pushes back.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#377

Earlier quoted context omitted.

It's overwhelmingly exceptionally rare, but famously SBF, Holmes, and Winterkorn.

Didn't they famously break actual laws though, not just "violating ethics"?

It's a bit reductive, but yes people are sent to prison for being convicted of crimes.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#378

Earlier quoted context omitted.

This is undefined. Without more information you don’t know the exact number of people. Riddle me this, why didn’t you do a better riddle?

Person 1: "I need chairs for two fathers and two sons to sit" Person 2: 'Okay, I have no idea how many chairs to grab, not enough information' - nobody ever (Person 2 has no ability to contribute to anything of economic value.)

Anyone who talks like person 1 contributes negative economic value.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#379

Earlier quoted context omitted.

Not quite. The idea that corporate employees are fundamentally "not average" and therefore more prone to unethical behaviour than the general population relies on a dispositional explanation (it's about the person's character). However, the vast majority of psychological research over the last 80 years heavily favours a situational explanation (it's about the environment/system). Everyone (in the field) got really in…

The Stanford prison experiment has been debunked many times : https://pubmed.ncbi.nlm.nih.gov/31380664/ - guards received instructions to be cruel from experimenters - guards were not told they were subjects while prisoners were - participants were not immersed in the simulation - experimenters lied about reports from subjects. Basically it is bad science and we can't conclude anything from it. I wouldn't rule out th…

- participants all self-selected into the study

They put an ad in a newspaper in San Francisco and then selected for apparent neurotypicality:

ZPE: https://en.wikipedia.org/wiki/Stanford_prison_experiment :

> Participants were recruited from the local community through an advertisement in the newspapers offering $15 per day ($119.41 in 2025) to male students who wanted to participate in a "psychological study of prison life".

Here's that newspaper ad: https://exhibits.stanford.edu/spe/catalog/cj859hr0956 :

> Steady P-Time Job

> [...]

> Male college students needed for psychological study of prison life. $15/day for 1-2 weeks beg Aug. 14. For further information & applications, come to Room 248, Jordan Hall, Stanford U.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#380

Earlier quoted context omitted.

it’s excel with extra steps. but for the linkedin layman, yes, it’s simplified version of brain neural networks.

Given this (even more linkedin layman) gross generalization, the human brain is not "excel with extra steps" how? Somehow the presense of chemicals and electrical signals and tissues makes the process not algorithmically reducible?

somehow the presence of signals doesn’t really equate intelligence. clearly
Post reply on HN