The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
31–40 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#32Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…
it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#33Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…
"How much money does it take to be rich and powerful like Anthropic intends?"
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#34I assume Anthropic will continue to tune the model, so I am not too bothered by this.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#35Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.
Our future is loonytoons.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#36Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#37I am using LLM to build some security tool, and I ran into this a few times. I have to come up with a reasoning to convince (?!!) Fable to continue the work without downgrading. I assume Anthropic will continue to tune the model, so I am not too bothered by this.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#38Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#39Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#40What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.