Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

111–120 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#111
So I suspect Anthropic started A/B testing or just plain testing this a while ago,

Tell HN: Claude flags biology / biotech questions https://news.ycombinator.com/item?id=47929885

Today, it's flagging population research questions,

    Using only the dataset you constructed, assess two questions:
     
    1. **Mortality:** do [GROUP] show mortality that differs
       from (a) your comparison groups and (b) era- and sex-matched US population
       expectations (e.g., SSA cohort life tables)?
    2. **Late-life outcomes:** define an endpoint you consider fair (justify it),
       and assess whether [GROUP] differs from comparators. State
       explicitly how your `documentation_depth` codings affect the strength of any
       conclusion — i.e., quantify or bound the ascertainment problem rather than waving at it.
    
    Choose your own methods and justify them. Report effect sizes with confidence intervals,
    not just p-values. State conclusions plainly, including "no detectable difference" if
    that is what your analysis shows — a null is an acceptable answer for either question
    independently. Document any additional judgment calls (index date for time-at-risk,
    reference population construction, endpoint definition) in the same decision-log style.
https://github.com/anthropics/claude-code/issues/66780

Censored because I'm writing a paper. :)

Oh and forget learning about chemistry. Only criminals want to learn organic chemistry. :(

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#113
post #95
post #64

I make privacy tooling and Fable 5 rejects the vast majority of my prompts to analyze and improve the software that I've written. It's bleak.

Why is this surprising or a problem?! It's a model demo, & their reasoning is reasonable and fair. Why all this drama.

Because you're being allowed to ask and work only on topics that a certain company decides.

Local inference has never been so important as it is now.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#114
post #9

Earlier quoted context omitted.

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…

That is for whatever it considers reverse-engineering the model to try to create a competing one.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#115
post #88

Earlier quoted context omitted.

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

Or if your "self-driving" system such as FSD / waymo slowed the car down once it detected you work in cybersecurity or at a rival automaker and you were attempting to reach the train station or the airport to make you miss a conference meetup.

Trains made by Newag were programmed to brick themselves if they detected a non-Newag workshop was repairing them.

https://news.ycombinator.com/item?id=38638865

https://news.ycombinator.com/item?id=38628635

https://news.ycombinator.com/item?id=38567687

https://news.ycombinator.com/item?id=38530885

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#116
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

> it won't just reject ML research, which I can understand I don't.

They don't want someone to piggyback Anthropic's Mythos to make their own Mythos with less effort than it cost Anthropic.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#117

So I suspect Anthropic started A/B testing or just plain testing this a while ago, Tell HN: Claude flags biology / biotech questions https://news.ycombinator.com/item?id=47929885 Today, it's flagging population research questions, Using only the dataset you constructed, assess two questions: 1. **Mortality:** do [GROUP] show mortality that differs from (a) your comparison groups and (b) era- and sex-matched US popula…

I was digging into some orbital mechanics questions and I assume it decided I was trying to backyard-science my way into an orbital-bombardment weapon. Kind of wild how this product's impression has gone from "wow, this is pretty neat" to "irreverent sack of dog shit you" in 24 hours almost solely on the back of a half-baked moderation system.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#118

Earlier quoted context omitted.

If the guardrails were so useless, people wouldn't be complaining about them.

People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.

> it's speed bump at best

To be fair, speed bumps work. If it's actually speed bumping nefarious activity, that gives authorities more time to react.

The correct place to police rogue nucleotides is at the labs. Not the compute layer.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#119
post #39

Fable is utterly useless with those guardrails for any serious it or life science work. Anthropic fucked me once a few months ago by closing down the subscription for any other harness, now it fucked me twice with buying again a subscription to find out their hyped model is unusable for normies. Using their products feels like a constant battle instead of a productive work day.. compare that with openai, not once did…

What do you mean that it closed your subscription for any other harness?

In any case that's what closed source (weights) for the masses means.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#120
post #2

DeepSeek is the only one that I can directly ask about vulnerabilities and it will give me a PoC. Although not as good as others, it has helped me with security research. The rest have guard rails that are so heavy, it makes them almost useless for cybersecurity.

Deepseek training is not finished yet, it's a preview.

And yes, it's an excellent model.

Post reply on HN