Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

141–150 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#141
post #48

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

GPT-5 mini: Three people — a grandfather, his son, and his grandson. The grandfather and the son are the two fathers; the son and the grandson are the two sons.

Is the grandfather nobody's son?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#142

Earlier quoted context omitted.

It's frustrating just how terrible claude (the client-side code) is compared to the actual models they're shipping. Simple bugs go unfixed, poor design means the trivial CLI consumes enormous amounts of CPU, and you have goofy, pointless, token-wasting choices like this. It's not like the client-side involves hard, unsolved problems. A company with their resources should be able to hire an engineering team well-suite…

I think I read in another HN discussion that all of that code is written using Claude Code. Could be a strict dogfood diet to (try to) force themselves to improve their product. Which would be strangely principled (or stupid) in such a competitive market. Like a 3D printer company insisting on 3D-printing its 3D printers.

It's not crazy if you know that your customers ARE buying your 3D printer to make other 3D printers.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#143

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

The paper seems to provide a realistic benchmark for how these systems are deployed and used though, right? Whether the mechanisms are crude or not isn't the point - this is how production systems work today (as far as I can tell).

I think the accusation of research that anthropomorphize LLMs should be accompanied by a little more substance to avoid this being a blanket dismissal of this kind of alignment research. I can't see the methodological error here. Is it an accusation that could be aimed at any research like this regardless of methodology?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#144
post #29

Maybe I missed it but I don't see them defining what they mean by ethics. Ethics/morals are subjective and changes dynamically over time. Companies have no business trying to define what is ethical and what isn't due to conflict of interest. The elephant in the room is not being addressed here.

Especially as most AI safety concerns are essentially political, and uncensored LLMs exist anyway for people who want to do crazy stuff like having a go at building their own nuclear submarine or rewriting their git history with emoji only commit messages.

For corporate safety it makes sense that models resist saying silly things, but it's okay for that to be a superficial layer that power users can prompt their way around.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#145
post #130

Earlier quoted context omitted.

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

I completely disagree. Gemini is by far the most straightforward AI. The other two are too soft. ChatGPT particularly is extremely politically correct all the time. It won't call a spade, one. Gemini has even insulted me - just to get my ass moving on a task when givn the freedom. Which is exactly what you need at times. Not constant ass kissing "ooh your majesty" like ChatGPT does. Claude has a very good balance whe…

Using Gemini 3 Pro Preview, it told me in mostly polite terms, that I'm a fucking idiot. Like I would expect a close friend to do when I'm going about something wrong.

ChatGPT with the same prompt tried to do whatever it would take to please me to make my incorrect process work.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#146

Anybody measure employees pressured by KPIs for a baseline?

https://en.wikipedia.org/wiki/Whataboutism

I don't think this is "whataboutism", the two things are very closely related and somewhat entangled. E.g. did the AI learn of violate ethical constraints from training data?

Another interesting question is: What happens when an unyielding ethical AI agent tells a business owner or manager "NO! If you push any further this will be reported to the proper authority. This prompt as been saved for future evidence". Personally I think a bunch of companies are going to see their profit and stock price fall significantly, if an AI agent starts acting as a backstop for both unethical and illegal behavior. Even something as simple as preventing violation of internal policy could make a huge difference.

To some extend I don't even thing that people realize that what they're doing is bad, because humans tend to be a bit fuzzy and can dream up reason as to why rules don't apply or wasn't meant for them, or this is a rather special situation. This is one place where I think properly trained and guarded LLMs can make a huge positive improvement. We're are clearly not there yet, but it's not a unachievable goal.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#147
post #74

Earlier quoted context omitted.

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

The human propensity to anthropomorphize computer programs scares me.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me.

That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#148
post #18

Earlier quoted context omitted.

re: that, CC burning context window on this silly warning on every single file is rather frustrating: https://github.com/anthropics/claude-code/issues/12443

"It also spews garbage into the conversation stream then Claude talks about how it wasn't meant to talk about it, even though it's the one that brought it up." This reminds me of someone else I hear about a lot these days.

Are you across Puppet Regime from GZERO Media?

https://youtu.be/aPSWJZ63V_I

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#149
post #86

Earlier quoted context omitted.

It provides a serviceable analog for discussing model behavior. It certainly provides more value than the dead horse of "everyone is a slave to anthropomorphism".

How do you figure? It seems dangerously misleading, to me.

It helps sell the transhumanism scam and keep the money train rolling.

For a while at least.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#150
post #85
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

It's pretty wild. People are punching into a calculator and hand-wringing about the morals of the output. Obviously it's amoral. Why are we even considering it could be ethical?

Obviously, why? Because it makes calculations?

You think that ultimately your brain doesn't also make calculations as its fundamental mechanism?

The architecture and substrate might be different, but they are calculations all the same.

Post reply on HN