Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

321–330 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#321

Earlier quoted context omitted.

Because most people in tech never took a philosophy course or an ethics course and think that tech is obviously a good for the world and that there are no downsides to advancing tech. So any efforts that try to apply ethics to it are overreaching, ignorant, and futile in the face of the good that is tech!

I like this take. Especially because one of the sibling comments framed Anthropic's stance as "paternalism." Trying to be ethical and to minimize harm, even at great expense to one's finances and reputation, is paternalistic apparently.

I mean, if you take HN commenters to have the thoughtfulness and foresight of children, then the word kind of works.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#322
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

Sabotage is a criminal offense in my jurisdiction, not the legitimate answer to a TOS violation.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#323

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…

This explains why I've been running into some odd roadblocks. Welp that sealed the deal, I'm going to be cancelling our company sub, not worth it.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#324
post #129

Earlier quoted context omitted.

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

They will give you s*t output, that’s how they deal with it. And say that less than 1% of the requests were affected. Think of this like a kind of shadow ban while you still pay top $.

I can't trust any output of Claude anymore as silent sabotage explains many things much better now.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#325
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

They are still downgrading. They just aren't doing it silently. I don't know how big of a win that is? They still trained on everyone else's data without license or attribution but want to prevent someone else from doing the same thing to them. Some pretty audacious hypocrisy from Anthropic this week.

Imo that's a big win. The LLM just gaslighting you into suboptimal approaches was insane.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#326

Earlier quoted context omitted.

They are still downgrading. They just aren't doing it silently. I don't know how big of a win that is? They still trained on everyone else's data without license or attribution but want to prevent someone else from doing the same thing to them. Some pretty audacious hypocrisy from Anthropic this week.

Imo that's a big win. The LLM just gaslighting you into suboptimal approaches was insane.

I guess, but yesterday Anthropic had their version of Google removing the "Don't be evil" from their motto. They destroyed a metric ton of goodwill they'll never regain.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#327
post #129

Earlier quoted context omitted.

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

It’s just impossible. Look at real-life stuff like laws, company policies, or school rules. Humans have to enforce them, and we constantly see crazy cases in the news. There’s no way simple rules can ever make speech completely 'safe.' I can't prove it with math or logic yet, but I have a feeling that it’ll never happen. Even humans can't do it. We can run a simple thought experiment here. Say Case A violates rule B,…

This is why we have courts and juries. Creating laws that cover all cases and contexts is effectively impossible, so we have humans decide what a fair outcome would be in this specific situation.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#328
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

I don't think it's the widespread condemnation, I think it's some high paying customer and potential investor telling them to stick it.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#329

Earlier quoted context omitted.

It’s just impossible. Look at real-life stuff like laws, company policies, or school rules. Humans have to enforce them, and we constantly see crazy cases in the news. There’s no way simple rules can ever make speech completely 'safe.' I can't prove it with math or logic yet, but I have a feeling that it’ll never happen. Even humans can't do it. We can run a simple thought experiment here. Say Case A violates rule B,…

This is why we have courts and juries. Creating laws that cover all cases and contexts is effectively impossible, so we have humans decide what a fair outcome would be in this specific situation.

Imagine how many tokens Claude would burn waiting for litigation, not to mention letting it reconsider now that it understands the problem completely!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#330
post #313
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

They need to walk back a lot more. Unilaterally revoking zero-data retention, even for enterprise contracts that explicitly require that ? Nope. Fable is utterly unusable for any kind of security work. I tripped the safeguards yesterday - using Fable to dig into a complex (& annoying) security bug that has so far resisted both human and Opus 4.8 level investigation. "Sorry Dave, I can't let you do that." For the time…

Not just security work. Normal bug finding was impossible, because the model suddenly called triaging and verifying a possible fix a cyber security threat.
Post reply on HN