Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

71–80 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#71
post #48

Earlier quoted context omitted.

Was this program available to independent security researchers or just established organizations? The docs you linked aren't very clear on this.

Any public research footprint seems to be enough, I applied as an individual and everyone I know who tried got accepted.

I have applied twice with half a dozen public CVEs and have been denied both times.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#72
post #9

Earlier quoted context omitted.

I've seen this claim a few times, but when I triggered the guardrails in Claude Code, it clearly notified me that it had switched to a different model ("something something for security purposes..."). Are you using Fable in Claude Code or in the browser?

They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.

> a LORA that's designed to inject bugs into your code

A statement like this, clearly, requires a reference.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#73

Earlier quoted context omitted.

it triggered for my.... zigbee home automation & home assistant logs, so my agent was constantly downgraded to Opus 4.8 even after I've changed it back. The false positives never stopped. "Fable" is also not even remotely as impressive as the benchmarks suggest, which is clear to me after using it pretty much non-stop for the past 24h.

It has to be sort of impressive, given that you tried so hard to use it instead of the regular Opus.

Considering that this is a brand new release of a frontier model that Anthropic is hyping hard, I'm not sure that the conclusion to draw from their repeated attempts to use it is that it's impressive... Anthropic is promising that it's impressive and we're all trying to test it out.

I, for one, have tried using it several times today and the guardrails kept switching the model back to Opus, so I have no clue if it's impressive or not.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#75
post #72

Earlier quoted context omitted.

They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.

> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.

From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#76
post #52

It's a marketplace. Someone else will outdo this inferior product.

OpenAI is the only real competition. Chinese models are 6-8 months behind Opus 4.8/GPT 5.5, and at least a year or more behind Mythos.

And it doesn't look like OpenAI will have a good answer to Mythos anytime soon. Based on what their chief scientist wrote to staff recently (https://archive.is/fN2pg), GPT 5.6 is a "meaningful improvement" over 5.5 - in other words, just a normal version bump. And no news or even rumors regarding GPT 6.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#77

I wonder how many millions they are wasting on putting up these guardrails when it's a completely useless exercise that is a speed bump at best.

If the guardrails were so useless, people wouldn't be complaining about them.

People are generally complaining about false positives. Now if you really wanna know what a real criminal organization would do... They'd just buy data center hardware even if it costs 200k because a successful targeted hit could yield far in excess of that. So yes it's speed bump at best.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#78
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

> it won't just reject ML research, which I can understand

I don't.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#79
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

The thing that I keep thinking about is the accounting / charging when it downgrades automatically.

Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price?

If the answer is no, could that be construed as fraud?

Post reply on HN