Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

251–260 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#251

I wonder how much of the violation of ethical, and often even legal constraints in the business world today one could tie not only to the KPI pressure but also to the the awful "better to ask for forgiveness than permission" mentality that is reinforced by many "leadership" books written up by burnt out mid-level veterans of Mideast wars, trying to make sense of their "careers" and pushing out their "learnings" on to…

>who during their "careers" in the military were in effect being "kept", by being provided housing, clothing and free meals. Long term I can see this happen for all humanity where AI takes over thinking and governance and humans just get to play pretend in their echo chambers. Might not even be a downgrade for current society.

    All Watched Over By Machines Of Loving Grace (Richard Brautigan)

    I like to think (and
    the sooner the better!)
    of a cybernetic meadow
    where mammals and computers
    live together in mutually
    programming harmony
    like pure water
    touching clear sky.

    I like to think
    (right now, please!)
    of a cybernetic forest
    filled with pines and electronics
    where deer stroll peacefully
    past computers
    as if they were flowers
    with spinning blossoms.

    I like to think
    (it has to be!)
    of a cybernetic ecology
    where we are free of our labors
    and joined back to nature,
    returned to our mammal
    brothers and sisters,
    and all watched over
    by machines of loving grace.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#252
post #206

Who defines "ethics"?

People and societies. Your question is an important one, but also one that has been extensively researched, documented and improved upon. Whole fields of science, like "Metaethics" deal with answering your question. Other fields of science with defining "normative ethics" aka ethics that "everyone agrees upon" and so on. I may have misread your question as a somewhat dismissive sarcastic take or as a "Ethics are nons…

Not quite. You are describing "kinds of ethics" after ethics is an already established concept. I.e. actual examples of human ethics. Now the question is who defines ethics as concept in general. Humans can have ethics, but is it applicable to the computer programs at all? Sure, programs can have programmed limitations, but is that called ethics at all? Does my Outlook client has ethics, only because it has configured rules? What is the difference between my email client automatically responding to an email with "salesforce" mentioned and an LLM program automatically responding to a query with the word "plutonium"?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#253

Earlier quoted context omitted.

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

Humans risk jail time, AIs not so much.

Do they, really? Which CEO went to jail for ethical violations?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#254
post #183

Earlier quoted context omitted.

A remarkable number of humans given really quite basic feedback will perform actions they know will very directly hurt or kill people. There are a lot of critiques about quite how to interpret the results but in this context it’s pretty clear lots of humans can be at least coerced into doing something extremely unethical. Start removing the harm one, two, three degrees and add personal incentives and is it that surpr…

> 2012, Australian psychologist Gina Perry investigated Milgram's data and writings and concluded that Milgram had manipulated the results, and that there was a "troubling mismatch between (published) descriptions of the experiment and evidence of what actually transpired." She wrote that "only half of the people who undertook the experiment fully believed it was real and of those, 66% disobeyed the experimenter".[29…

[flagged]

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#256
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me. That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

it’s excel with extra steps. but for the linkedin layman, yes, it’s simplified version of brain neural networks.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#257

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

From an IBM training manual (1979): >A computer can never be held accountable >Therefore a computer must never make a management decision The (EDITED) corollary would arguably be: >Corporations are amoral entities which are potentially immortal who cannot be placed behind bars. Therefore they should never be given the rights of human beings. (potentially, not absolutely immortal --- would wording as "not mortal by es…

I changed my stance on "immoral" corporations:

Legal systems are the ones being "immoral" and "unethical" and "not just", not "righteous", not fair. They represent entire nations and populations while corpos represent interests of subsets of customers and "sponsors".

If corpos are forced to pivot because they are behaving ugly, they will ... otherwise they might lose money (although that is barely an issue anymore, given how you can offset almost any kind of loss via various stock market schemes).

But the entire chain upstream of law enforcement behaves ugly and weak, which is the fault of humanities finest and best earning "engineers".

Just take a sabbatical and fix some of that stuff ...

>> I mean you and your global networks got money and you can even stay undetected, so what the hell is the issue? Personal preference? Damn it, I guess that settles that. <<

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#258

Earlier quoted context omitted.

It provides a serviceable analog for discussing model behavior. It certainly provides more value than the dead horse of "everyone is a slave to anthropomorphism".

Where is Pratchett when we need him? I wonder how he would have chose to anthropomorphize anthropomorphism. A sort of meta anthropomorphization.

Maybe a being/creature that looked like a person when you concentrated on it and then was easily mistaken as something else when you weren't concentrating on it.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#259
post #67

Earlier quoted context omitted.

That is not a meaningful benchmark. They just made shit up. Regardless of whether any company cares or not, the whole concept of "AI safety" is so silly. I can't believe anyone takes it seriously.

Would you mind explaining your point a view? Or point me to ressources making you think so?

What can be asserted without evidence can also be dismissed without evidence. The benchmark creators haven't demonstrated that higher scores result in fewer humans dying or any meaningful outcome like that. If the LLM outputs some naughty words that's not an actual safety problem.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#260
post #5

https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range

Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.

The fact that the guy leading the development of Gemini was on Epstein's island is probably unrelated.
Post reply on HN