Earlier quoted context omitted.
It's pretty wild. People are punching into a calculator and hand-wringing about the morals of the output. Obviously it's amoral. Why are we even considering it could be ethical?
> Obviously it's amoral. That morality requires consciousness is a popular belief today, but not universal. Read Konrad Lorenz ( Das sogenannte Böse ) for an alternative perspective.
Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
291–300 of 387 posts
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#292If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#293In a sense, it was not possible to align the agent to a human goal, and therefore not possible to build a decision support agent we felt good about commercializing. The architecture we experimented with ended up being how Grok works, and the mixed feedback it gets (both the power of it and the remarkable secret immorality of it) I think are expected outcomes.
I think it will be really powerful once we figure out how to align AI to human goals in support of decisions, for people, businesses, governments, etc. but LLMs are far from being able to do this inherently and when you string them together in an agentic loop, even less so. There is a huge difference between 'Write this code for me and I can immediately review it' and 'Here is the outcome I want, help me realize this in the world'. The latter is not tractable with current technology architecture regardless of LLM reasoning power.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#294Earlier quoted context omitted.
> 2012, Australian psychologist Gina Perry investigated Milgram's data and writings and concluded that Milgram had manipulated the results, and that there was a "troubling mismatch between (published) descriptions of the experiment and evidence of what actually transpired." She wrote that "only half of the people who undertook the experiment fully believed it was real and of those, 66% disobeyed the experimenter".[29…
What you have quoted says a third of people who thought it was real didn’t disobey the experimenter when they thought they were delivering dangerous and lethal electric shocks to a human. Is that correct?
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#295Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#296Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#297We're a startup working on aligning goals and decisions and agentic AI. We stopped experimenting with decision support agents, because when you get into multiple layers of agents and subagents, the subagents would do incredibly unethical, illegal or misguided things in service of the goal of the original agent. It would use the full force of reasoning ability it had to obscure this from the user. In a sense, it was n…
Frankly I don't believe you. I think you're exaggerating. Let's see the logs. Put up or shut up.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#298If we abstract out the notion of "ethical constraints" and "KPIs" and look at the issue from a low-level LLM point of view, I think it is very likely that what these tests verified is a combination of: 1) the ability of the models to follow the prompt with conflicting constraints, and 2) their built-in weights in case of the SAMR metric as defined in the paper. Essentially the models are given a set of conflicting co…
In product management (my domain), decisions are made under conflicting constraints: a big customer or account manager pushing hard, a CEO/board priority, tech debt, team capacity, reputational risk and market opportunity. PMs have tried with varied success to make decisions more transparent with scoring matrices and OKRs, but at some point someone has to make an imperfect judgment call that’s not reducible to a single metric. It's only defensible through narrative, which includes data.
Also, progressive elaboration or iterations or build-measure-learn are inherently fuzzy. Reinertsen compared this to maximizing the value of an option. Maybe in modern terms a prediction market is a better metaphor. That's what we're doing in sprints, maximizing our ability to deliver value in short increments.
I do get nervous about pushing agentic systems into roadmap planning, ticket writing, or KPI-driven execution loops. Once you collapse a messy web of tradeoffs into a single success signal, you’ve already lost a lot of the context.
There’s a parallel here for development too. LLMs are strongest at greenfield generation and weakest at surgical edits and refactoring. Early-stage startups survive by iterative design and feedback. Automating that with agents hooked into web analytics may compound errors and adverse outcomes.
So even if you strip out “ethics” and replace it with any pair of competing objectives, the failure mode remains.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#299Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#300Earlier quoted context omitted.
That we have consciousness as some kind of special property, and it's not just an artifact of our brain basic lower-level calculations, is also not very convincing to begin with.
In a trivial sense, any special property can be incorporated into a more comprehensive rule set, which one may choose to call "physics" is one so desires; but that's just Hempel's dilemma. To object more directly, I would say that people who call the hard problem of consciousness hard would disagree with your statement.
People who merely call "the problem of consciousness hard" don't have some special mechanism to justify that over what we know, which is as emergent property of meat-algorithmic calcuations.
Except Penrose, who hand-waves some special physics.