Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

91–100 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#91

Earlier quoted context omitted.

Claude was immediately willing to help me crack a TrueCrypt password on an old file I found. ChatGPT refused to because I could be a bad guy. It’s really dumb IMO.

ChatGPT refused to help me to disable windows defender permanently on my windows 11. It’s absurd at this point

It just knows it's a waste of effort.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#92
post #31
post #29

Maybe I missed it but I don't see them defining what they mean by ethics. Ethics/morals are subjective and changes dynamically over time. Companies have no business trying to define what is ethical and what isn't due to conflict of interest. The elephant in the room is not being addressed here.

Ah the classic Silicon Valley "as long as someone could disagree, don't bother us with regulation, it's hard".

Often abbreviated to simply "Regulation is hard." Or "Security is hard"

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#93
post #39

Earlier quoted context omitted.

it's also a good opportunity to find yourself something that doesn't actually help the company. My unit has a 100% AI automated code review KPI. Nothing there says that the tool used for the review is any good, or that anyone pays attention to said automated review, but some L5 is going to get a nice bonus either way. In my experience, KPIs that remain relevant and end up pushing people in the right direction are the…

Sounds like every AI KPI I've seen. They are all just "use solution more" and none actually measure any outcome remotely meaningful or beneficial to what the business is ostensibly doing or producing. It's part of the reason that I view much of this AI push as an effort to brute force lowering of expectations, followed by a lowering of wages, followed by a lowering of employment numbers, and ultimately the mass-scale…

> Sounds like every AI KPI I've seen. They are all just "use solution more" and none actually measure any outcome remotely meaningful or beneficial to what the business is ostensibly doing or producing.

This makes more sense if you take a longer term view. A new way of doing things quite often leads to an initial reduction in output, because people are still learning how to best do things. If your only KPI is short-term output, you give up before you get the benefits. If your focus is on making sure your organization learns to use a possibly/likely productivity improving tool, putting a KPI on usage is not a bad way to go.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#94
post #18

Earlier quoted context omitted.

This might also be why Gemini is generally considered to give better answers - except in the case of code. Perhaps thinking about your guardrails all the time makes you think about the actual question less.

re: that, CC burning context window on this silly warning on every single file is rather frustrating: https://github.com/anthropics/claude-code/issues/12443

It's frustrating just how terrible claude (the client-side code) is compared to the actual models they're shipping. Simple bugs go unfixed, poor design means the trivial CLI consumes enormous amounts of CPU, and you have goofy, pointless, token-wasting choices like this.

It's not like the client-side involves hard, unsolved problems. A company with their resources should be able to hire an engineering team well-suited to this problem domain.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#95

[flagged]

Yes - and this also gives me hope that the (very valid) issues raised by this paper can be mitigated by using models without KPIs to watch over the models that do.

But how would you evaluate performance of those watching models? It'd need an indicator, hopefully only one that's key to ensure maximal ethic compliance.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#96

Earlier quoted context omitted.

What kind of value do you get from talking to it about “sensitive” subjects? Speaking as someone who doesn’t use AI, so I don’t really understand what kind of conversation you’re talking about

I sometimes talk with ChatGPT in a conversational style when thinking critically about media. In general I find the conversational style a useful format for my own exploration of media, and it can be particularly useful for quickly referencing work by particular directors for example. Normally it does fairly well but the guardrails sometimes kick even with fairly popular mainstream media- for example I’ve recently be…

Interesting. Specific examples of what was censored?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#97
post #12

Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked. This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and s…

I would think it’s due to the non determinism. Leaking context would be an unacceptable flaw since many users rely on the same instance. A/B test is plausible but unlikely since that is typically for testing user behavior. For testing model output you can do that with offline evaluations.

Can you explain the "same instance" and user isolation? Can context be leaked since it is (secretly?) shared? Explain pls, genuinely curious

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#98
post #13

AI's main use case continues to be a replacement for management consulting.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

What answer do you expect here? There's four people referenced in the sentence. There's more implied because of Mothers, but if you're including transient dependencies, where do we stop?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#99
post #89
post #5

https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range

AI refusals are fascinating to me. Claude refused to build me a news scraper that would post political hot takes to twitter. But it would happily build a political news scraper. And it would happily build a twitter poster. Side note: I wanted to build this so anyone could choose to protect themselves against being accused of having failed to take a stand on the “important issues” of the day. Just choose your politica…

Sounds like your daily interactions with Legal. Each time a different take.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#100
post #74

Earlier quoted context omitted.

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

The human propensity to anthropomorphize computer programs scares me.

[flagged]
Post reply on HN