Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

11–20 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#11
post #7

The bio angle is crazy to think about - imagine a health crisis triggered by LLM. What a time we live in.

This is all so amazing and good. These are exciting times we’re living in. Can’t wait to see what the future holds.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#12
post #2

DeepSeek is the only one that I can directly ask about vulnerabilities and it will give me a PoC. Although not as good as others, it has helped me with security research. The rest have guard rails that are so heavy, it makes them almost useless for cybersecurity.

[deleted]

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#13
post #9
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

Different restrictions. ML gets treated differently from the rest.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#14
post #9
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

It's from the model card:

> unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT).

https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

(stolen from https://jonready.com/blog/posts/claude-fable5-is-allowed-to-...)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#16
post #9
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

Specifically only ML research

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#17

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

If they can train the classifier to have fewer false positives that would be great.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#18
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

"How much money does it take to be rich and powerful like Anthropic intends?"

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#19
It seems like they've given up on the idea of the Cyber Verification Program https://support.claude.com/en/articles/14604842-real-time-cy...

When Opus 4.7 was introduced it started refusing anything cyber-adjacent (as an API error message, not a conversational refusal), until you applied for CVP, which made it more sensible again.

In Opus 4.8 it doesn't seem to help much, you just get refusals as prose rather than API errors. And now in Fable you don't get anything at all.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#20
post #15

Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.

I've done this, including the hardcoded refusal strings that already exist in claude code. It won't stop a real attacker, but I still find it really funny when you're trying to use one of the AI tools and it gives you a random refusal and you don't know why, wastes a little bit of time.
Post reply on HN