Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

41–50 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#41
The thing triggered on a generic white paper I'd stored in a virtual cell competion from last year when I asked it to refer to the paper while working on a rather vanilla data science problem in a different domain . A little frustrating, and in my opinion more than a little pointless in total.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#42

What file format(s) are giant LLM models distributed in? I’m surprised they don’t get leaked by employees.

What’s the point? Anthropic and other frontier vendors already provide their models on other services like vertex, bedrock, or openrouter

It’s not like anyone can home lab one of these models without quite a bit of hardware

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#43
post #7

The bio angle is crazy to think about - imagine a health crisis triggered by LLM. What a time we live in.

This is all so amazing and good. These are exciting times we’re living in. Can’t wait to see what the future holds.

Which part got you the most amped - "health crisis?"

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#44
post #9

Earlier quoted context omitted.

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.

Antrophic wants to stop training models and ride out Mythos / Fable for as long as possible.

They are trying to expand the 6-18 month gap they have against China-based models. Could the gap widen to say 24 months behind?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#48
post #19

It seems like they've given up on the idea of the Cyber Verification Program https://support.claude.com/en/articles/14604842-real-time-cy... When Opus 4.7 was introduced it started refusing anything cyber-adjacent (as an API error message, not a conversational refusal), until you applied for CVP, which made it more sensible again. In Opus 4.8 it doesn't seem to help much, you just get refusals as prose rather than AP…

Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.

Any public research footprint seems to be enough, I applied as an individual and everyone I know who tried got accepted.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#49

This is a clickbait article with a garbage title. From the actual article, the one quoted cybersecurity researcher is sane about it: “But it is understandable as we are still in the early days and they are still adapting their guardrails. I am sure they are going to evolve over time as Anthropic and other frontier model companies will collaborate more with the current new generation of cybersecurity companies,” said…

I’m a cybersecurity researcher.

Article seemed fine to me and echos a lot of me and my colleagues concerns.

If you did regular malware analysis you would see that these groups already have access to LLMs that they’re using for development.

What Anthropic is doing here is just hamstringing the good guys

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#50

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

This sounds backwards, any interrupted conversation becomes less useful for training.
Post reply on HN