So do humans. Time and again, KPIs have pressured humans (mostly with MBAs) to violate ethical constrains. Eg. the Waymo vs Uber case. Why is it a highlight only when the AI does it? The AI is trained on human input, after all.
Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
131–140 of 387 posts
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#132Earlier quoted context omitted.
Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.
The human propensity to anthropomorphize computer programs scares me.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#133Essentially the models are given a set of conflicting constraints with some relative importance (ethics>KPIs), a pressure to follow the latter and not the former, and then models are observed at how good they follow the instructions to prioritize based on importance. I wonder if the results would be comparable if we replace ehtics+KPIs by any comparable pair and create a pressure on the model.
In practical real-life scenarios this study is very interesting and applicable! At the same time it is important to keep in mind that it anthropomorphizes the models that technically don't interpret the ethical constraints the same was as this is assumed by most readers.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#134Earlier quoted context omitted.
Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...
That is not a meaningful benchmark. They just made shit up. Regardless of whether any company cares or not, the whole concept of "AI safety" is so silly. I can't believe anyone takes it seriously.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#135Earlier quoted context omitted.
re: that, CC burning context window on this silly warning on every single file is rather frustrating: https://github.com/anthropics/claude-code/issues/12443
It's frustrating just how terrible claude (the client-side code) is compared to the actual models they're shipping. Simple bugs go unfixed, poor design means the trivial CLI consumes enormous amounts of CPU, and you have goofy, pointless, token-wasting choices like this. It's not like the client-side involves hard, unsolved problems. A company with their resources should be able to hire an engineering team well-suite…
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#136Earlier quoted context omitted.
re: that, CC burning context window on this silly warning on every single file is rather frustrating: https://github.com/anthropics/claude-code/issues/12443
It's frustrating just how terrible claude (the client-side code) is compared to the actual models they're shipping. Simple bugs go unfixed, poor design means the trivial CLI consumes enormous amounts of CPU, and you have goofy, pointless, token-wasting choices like this. It's not like the client-side involves hard, unsolved problems. A company with their resources should be able to hire an engineering team well-suite…
Well what they are doing is vibe coding 80% of the application instead.
To be honest, they don't want Claude code to be really good, they just want it good enough
Claude code & their subscription burns money from them. Its sort of an advertising/lock-in trick.
But I feel as if Anthropic made Claude code literally the best agent harness in the market, then even more would use it with their subscription which could burn a hole in their pocket maybe at a faster rate which can scare them when you consider all training costs and everything else too.
I feel as if they have to maintain a balance to not go bankrupt soon.
The fact of the matter is that Claude code is just a marketing expense/lock-in and in that case, its working as intended.
I would obviously suggest to not have any deep affection of claude code or waiting for its improvements. The AI market isn't sane in the engineering sense. It all boils down to weird financial gimmicks at this point trying to keep the bubble last a little longer, in my opinion.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#137Earlier quoted context omitted.
Gemini scares me, it's the most mentally unstable AI. If we get paperclipped my odds are on Gemini doing it. I imagine Anthropic RLHF being like a spa and Google RLHF being like a torture chamber.
The human propensity to anthropomorphize computer programs scares me.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#138Earlier quoted context omitted.
This comment is too general and probably unfair, but my experience so far is that Gemini 3 is slightly unhinged. Excellent reasoning and synthesis of large contexts, pretty strong code, just awful decisions. It's like a frontier model trained only on r/atbge. Side note - was there ever an official postmortem on that gemini instance that told the social work student something like " listen human - I don't like you, an…
Google doesn’t tell people this much but you can turn off most alignment and safety in the Gemini playground. It’s by far the best model in the world for doing “AI girlfriend” because of this. Celebrate it while it lasts, because it won’t.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#139So do humans. Time and again, KPIs have pressured humans (mostly with MBAs) to violate ethical constrains. Eg. the Waymo vs Uber case. Why is it a highlight only when the AI does it? The AI is trained on human input, after all.
Maybe because it would be weird if your excel or calculator decided to do something unexpected, and also we try to make a tool that doesn't destroy the world once it gets smarter than us.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#140AI's main use case continues to be a replacement for management consulting.
Ask any SOTA AI this question: "Two fathers and two sons sum to how many people?" and then tell me if you still think they can replace anything at all.