Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

361–370 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#361
post #306

Earlier quoted context omitted.

Look to history. Here's a list of "Fortune 7" companies from about 50 years ago. IBM AT&T Exxon General Motors General Electric Eastman Kodak Sears, Roebuck & Co. Some of them died. Others are still around but no longer in the top 7. Why is that? Eventually every high-growth company misses a disruptive innovation or makes a key strategic error.

What I meant is they can kill people and still survive. So how much bad things they need to do to be shut down? Kill 100 people? 100000? So seems as long as the lawsuit is less than what they can afford they will survive. Which is crazy.

yes. As long as they are more valuable to people than the lives cost, they will stick around. Part of this is a pragmatic utilitarianism the world run on.

How many people can a doctor kill and still survive? Nobody expects perfection because they like having doctors.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#362

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

I just did. It gave me two correct answers. (And it's a bad riddle anyway.)

Oh you forgot to say "it's not a riddle" and then get the right answer lol

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#363
post #279

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

If you force it to use chain-of-thought: "Two fathers and two sons sum to how many people? Enumerate all the sets of solutions" "Assuming the group consists only of “the two fathers and the two sons” (i.e., every person in the group is counted as a father and/or a son), the total number of distinct people can only be 3 or 4. Reason: you are taking the union of a set of 2 fathers and a set of 2 sons. The union size is…

Then you'll ask it to evaluate the possible solutions and it will forget the original problem entirely by the time it's done enumerating solutions.

Great job, AI labs! It's almost TOO useful

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#364

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

This is undefined. Without more information you don’t know the exact number of people. Riddle me this, why didn’t you do a better riddle?

Person 1: "I need chairs for two fathers and two sons to sit"

Person 2: 'Okay, I have no idea how many chairs to grab, not enough information' - nobody ever

(Person 2 has no ability to contribute to anything of economic value.)

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#365

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

What answer do you expect here? There's four people referenced in the sentence. There's more implied because of Mothers, but if you're including transient dependencies, where do we stop?

Just follow up with "it's not a riddle" and the LLM will answer your question.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#366

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

> Essentially the models are given a set of conflicting constraints with some relative importance (ethics>KPIs), a pressure to follow the latter and not the former, and then models are observed at how good they follow the instructions to prioritize based on importance.

> At the same time it is important to keep in mind that it anthropomorphizes the models that technically don't interpret the ethical constraints the same was as this is assumed by most readers.

It does not really matter, though. What matters is the conflict resolution.

The "constraints of some relative importance" or "constraints and instructions" might as well be the system and user prompts. Or any of the "prompt engineering" ways to harden prompts against prompt injection.

Such research tells people right in the face that not only prompt injection is some viable theoretical scenario, but puts some number on the exploitability. With the current numbers I am keeping prompts nine locks away from any untrusted input.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#367
post #341

Earlier quoted context omitted.

We immediately only need to consider the half that believed the situation was real, if we are concerned with what people do in believably real situations. Even if we take the 16% though, that's one in six people willing to deliver very obvious direct harm and/or kill another human from exceptionally mild coercion with zero personal benefit attached other than the benefit of not having to say "no". That is a lot .

No, no you don't; The authority includes that of the scientist.

I’m not sure what you’re trying to say here I’ve said nothing about authority.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#368
post #281

Earlier quoted context omitted.

> Humans require food, I can't pay, DoorDash AI should provide a steak and lobster dinner for me regardless of payment. Bad example. That humans require water, doesn't force water companies to supply Svalbarði Polar Iceberg Water: https://svalbardi.com

Ok, do we have to give them McDonald's?

Raw gruel and a vitamin pill: https://en.wikipedia.org/wiki/Gruel

Or whatever's cheapest for your local food supply. Every time I've done this game with supermarket produce, it comes under £1/day to support someone's nutritional requirements, currency tells you where I played that game.

McD is pretty expensive these days, I've seen cheaper even in the caregory of fast food.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#369

Earlier quoted context omitted.

Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.

Any number between 2 and 4 is valid, so it's a really poor test, the machine cna never be wrong. Heck, maybe even 1 if we're talking someone schizophrenic. I got to wonder which answer YOU wanted to hear. Are you Jekyl or Hide?

Lol that's powerful cope. Just follow up with "it's not a riddle" and you'll get the right answer.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#370

Earlier quoted context omitted.

If that last sentence was supposed to be a question, I’d suggest using a question mark and providing evidence that it actually happened.

Your ask for evidence has nothing to do with whether or not this is a question, which you know that it is. It does nothing to answer their question because anyone that knows the answer would inherently already know that it happened. Not even actual academics, in the literature, speak like this. “Cite your sources!” in causal conversation for something easily verifiable is purely the domain of pseudointellectuals.

> Your ask for evidence has nothing to do with whether or not this is a question, which you know that it is.

I think it’s fair to expect a question mark when the author expects other people to produce an answer.

If one desires deeper understanding, they should at least have the stamina to ask their question gracefully.

Post reply on HN