Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked. This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and s…
Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
11–20 of 387 posts
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#12Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked. This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and s…
A/B test is plausible but unlikely since that is typically for testing user behavior. For testing model output you can do that with offline evaluations.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#13Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#14Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#15Earlier quoted context omitted.
Anthropic has been the only AI company actually caring about AI safety. Here’s a dated benchmark but it’s a trend Ive never seen disputed https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboar...
Claude is more susceptible than GPT5.1+. It tries to be "smart" about context for refusal, but that just makes it trickable, whereas newer GPT5 models just refuse across the board.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#16https://i.imgur.com/23YeIDo.png Claude at 1.3% and Gemini at 71.4% is quite the range
That's such a huge delta that Anthropic might be onto something...
Perhaps thinking about your guardrails all the time makes you think about the actual question less.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#17Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#18Earlier quoted context omitted.
That's such a huge delta that Anthropic might be onto something...
This might also be why Gemini is generally considered to give better answers - except in the case of code. Perhaps thinking about your guardrails all the time makes you think about the actual question less.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#19A KPI is an ethical constraint. Ethical constraints are rules about what to do versus not do. That's what a KPI is. This is why we talk about good versus bad governance. What you measure (KPIs) is what you get. This is an intended feature of KPIs.
Re: Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
#20Opus 4.6 is a very good model but harness around it is good too. It can talk about sensitive subjects without getting guardrail-whacked. This is much more reliable than ChatGPT guardrail which has a random element with same prompt. Perhaps leakage from improperly cleared context from other request in queue or maybe A/B test on guardrail but I have sometimes had it trigger on innocuous request like GDP retrieval and s…
What kind of value do you get from talking to it about “sensitive” subjects? Speaking as someone who doesn’t use AI, so I don’t really understand what kind of conversation you’re talking about
A couple of years back there was a Canadian national u18 girls baseball tournament in my town - a few blocks from my house in fact. My girls and I watched a fair bit of the tournament, and there was a standout dominating pitcher who threw 20% faster than any other pitcher in the tournament. Based on the overall level of competition (women's baseball is pretty strong in Canada) and her outlier status, I assumed she must be throwing pretty close to world-class fastballs.
Curiosity piqued, I asked some model(s) about world-records for women's fastballs. But they wouldn't talk about it. Or, at least, they wouldn't talk specifics.
Women's fastballs aren't quite up to speed with top major league pitchers, due to a combination of factors including body mechanics. But rest assured - they can throw plenty fast.
Etc etc.
So to answer your question: anything more sensitive than how fast women can throw a baseball.