Live data from Hacker News

Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

arxiv.org

311–320 of 387 posts

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#311
post #288

Earlier quoted context omitted.

It would also be interesting to see how humans perform on the same kind of tests. Violating ethics to improve KPI sounds like your average fortune 500 business.

So, I kind of get this sentiment. There is a lot of goal post moving going on. "The AIs will never do this." "Hey they're doing that thing." "Well, they'll never do this other thing." Ultimately I suspect that we've not really thought that hard about what cognition and problem solving actually are. Perhaps it's because when we do we see that the hyper majority of our time is just taking up space with little pockets o…

At least it is possible for an unethical person to face meaningful consequences and change their behavior.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#312

Earlier quoted context omitted.

How is a corporation "immortal"? What is the oldest corporation in the world? I mean, aside from churches and stuff. Corporations can die or be killed in numerous ways. Not many of them will live forever. Most will barely outlive a normal human's lifespan. By definition, since a corporation comprises a group of people, it could never outlive the members, should they all die at some point. Let us also draw a distincti…

What needs to do a company from fortune 7 to die? If kills 1 person they won’t close Google. If steals 1 billion, won’t close either. So what needs to do such a company to be closed down? I think it’s almost impossible to shut down

It took an armed rebellion and two acts of parliament to kill the British East India Company.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#314

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

At the very least it shows the capability of the current restrictions are deeply lacking and can be easily thwarted.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#315

If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…

If you want absolute adherence to a hierarchy of rules you'll quickly find it difficult - see I,Robot by Asimov for example. An LLM doesn't even apply rules, it just proceeds with weights and probabilities. To be honest, I think most people do this too.

You're using fiction writing as an example?

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#316
post #230

Earlier quoted context omitted.

> 2012, Australian psychologist Gina Perry investigated Milgram's data and writings and concluded that Milgram had manipulated the results, and that there was a "troubling mismatch between (published) descriptions of the experiment and evidence of what actually transpired." She wrote that "only half of the people who undertook the experiment fully believed it was real and of those, 66% disobeyed the experimenter".[29…

Milgram was flawed, sure. However, you can look at videos of ICE agents being surprised that their community think they're evil and doing evil, when they think they're just law enforcement. There was not even a need for coercion there, only story-telling.

Incorrect. ICE is built off the background of 30-50 years of propaganda against "immigrants", most of it completely untrue.

The same is done for "benefits scroungers", despite the evidence being that welfare fraud only accounts for approximately 1-5% of the cost of administering state welfare, and state welfare would be about 50%+ cheaper to administer if it was a UBI rather than being means-tested. In fact, much of the measures that are implemented with the excuse of "we need to stop benefits scroungers", such as testing if someone is disabled enough to work or not, etc. are simulatenously ineffective and make up most of the cost.

Nevertheless, "benefits scroungers" has entered the zeitgeist in the UK (and the US) because of this propaganda.

The same is true for propaganda against people who have migrated to the UK/US. Many have done so as asylum seekers under horrifying circumstances, and many die in the journey. However, instead of empathy, the media greets them with distaste and horror — dehumanising them in a fundamentally racist way, specifically so that a movement that grants them rights as a workforce never takes off, so that companies can employ them for zero-hour contracts to do work in conditions that are subhuman, and pay them substantially less than minimum wage (It's incredibly beneficial for the economy, unfortunately).

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#317

Earlier quoted context omitted.

Humans risk jail time, AIs not so much.

Do they, really? Which CEO went to jail for ethical violations?

Yeah, it’s exceptionally rare for CEOs, but they’re not the only one’s behaving unethically at work. There’s often a scapegoat.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#318
post #74

Earlier quoted context omitted.

The human propensity to anthropomorphize computer programs scares me.

The human propensity to call out as "anthropomorphizing" the attributing of human-like behavior to programs built on a simplified version of brain neural networks, that train on a corpus of nearly everything humans expressed in writing, and that can pass the Turing test with flying colors, scares me. That's exaxtly the kind of thing that makes absolute sense to anthropomorphize. We're not talking about Excel here.

> programs built on a simplified version of brain neural networks

Not even close. "Neural networks" in code are nothing like real neurons in real biology. "Neural networks" is a marketing term. Treating them as "doing the same thing" as real biological neurons is a huge error

>that train on a corpus of nearly everything humans expressed in writing

It's significantly more limited than that.

>and that can pass the Turing test with flying colors, scares me

The "turing test" doesn't exist. Turing talked about a thought experiment in the very early days of "artificial minds". It is not a real experiment. The "turing test" as laypeople often refer to it is passed by IRC bots, and I don't even mean markov chain based bots. The actual concept described by Turing is more complicated than just "A human can't tell it's a robot", and has never been respected as an actual "Test" because it's so flawed and unrigorous.

Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs

#320
post #85

Earlier quoted context omitted.

It's pretty wild. People are punching into a calculator and hand-wringing about the morals of the output. Obviously it's amoral. Why are we even considering it could be ethical?

Obviously, why? Because it makes calculations? You think that ultimately your brain doesn't also make calculations as its fundamental mechanism? The architecture and substrate might be different, but they are calculations all the same.

Brains do not "make calculations". Biological neurons do not "make calculations"

What they do is well described by a bunch of math. You've got the direction of the arrow backwards. Map, territory, etc.

Post reply on HN