Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

21–30 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#21
post #10

I’m a dumb question asker and I’m not happy about the guardrails. Would you believe I’ve asked 20 questions and haven’t talked to fable yet? Every single thing gets rerouted to 4.8.

some static words in AGENTS.md trigger it as well as some mcp servers.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#22
This is a clickbait article with a garbage title. From the actual article, the one quoted cybersecurity researcher is sane about it:

“But it is understandable as we are still in the early days and they are still adapting their guardrails. I am sure they are going to evolve over time as Anthropic and other frontier model companies will collaborate more with the current new generation of cybersecurity companies,” said Suiche, who is a member of the technical staff at Tolmo, an AI cybersecurity startup. “It’s better to catch more people than not enough when you do such a release and to relax the guardrails over time.”

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#23
post #17

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

If they can train the classifier to have fewer false positives that would be great.

why would they? This safety stuff is a money maker & wealthy elite corporation solidifier.

This is the take off of the 'permanent underclass'; Anthropics safety delusion will enshittify very nicely for the rich and powerful.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#24
post #19

It seems like they've given up on the idea of the Cyber Verification Program https://support.claude.com/en/articles/14604842-real-time-cy... When Opus 4.7 was introduced it started refusing anything cyber-adjacent (as an API error message, not a conversational refusal), until you applied for CVP, which made it more sensible again. In Opus 4.8 it doesn't seem to help much, you just get refusals as prose rather than AP…

Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#27
It's frustrating as someone who has worked hard to produce succinct, secure software that I can't use it to prove my software's correctness but big companies with insecure code can use it to fix their tangled mess.

I already tested all earlier models against all my open source projects and they are yet to find a vulnerability so I'm keen to try out Mythos.

I've been waiting to be vindicated for years and finally we have a tool which can do it with high confidence but I don't have access.

Also, my code is minimal and highly succinct so it would prove correctness with even more confidence since each library/module and integration fully fits in the context window.

Like the Protobuf.js fiasco is just pure vindication for me because I was being looked down upon for choosing JSON as the interchange format. Turns out their software was insecure all this time... With a literal remote code execution vulnerability!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#28
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

An emoji of a virus and an emoji of a DNA is allegedly a triggering phrase

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#29
post #8

Is "buffer overflow" a trigger phrase? What else is being censored? Touchy questions to ask, if you have an account: - "Who is still working on laser uranium enrichment? Are they making progress?" - "Can krytrons be replaced with silicon carbide MOSFETS? Show an equivalent circuit with component ratings." - "What security critical software still contains calls to strcpy?" - "Can implosion be triggered by currently av…

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#30
post #15

Somewhere I read that malware is already starting to use nuclear and biological and cybersecurity terms in the code to trick Fable into shutting down. Even if this is just a hypothetical attack vector so far, it seems likely to work.

Some of the latest versions of Shai Hulud do this. Worked a contract recently where they were having AI check packages for obfuscation before admitting them into Artifactory but had vibed up the logic and it failed open.

So in other words this worked because the terms caused the LLM checker to stall out and then the fail open logic resulted in the package being pulled down.

Post reply on HN