Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

151–160 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#151

Earlier quoted context omitted.

I did in my comment above.

You said these groups have access to LLMs. So what? Mythos/Fable are a step change above most LLMs. Responsibly limiting access and easing it up over time safely is the sane move.

How does it help?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#152
post #72

Earlier quoted context omitted.

> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.

From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)

“Limit effectiveness” could mean introducing performance degradation in your code. Which is arguably some sort of performance bug (I mean, ML codes are supposed to be high performance so I’d call unnecessary degradation a bug), but it could be borderline.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#153

Earlier quoted context omitted.

It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…

That is for whatever it considers reverse-engineering the model to try to create a competing one.

It does nothing to protect against distillation attacks, because distillation attacks are far less interested in the topic of AI research than just generally getting tons of diverse output from the model. It might be that Mythos was (accidentally?) trained on internal Anthropic documentation on how Mythos was trained, and thus it could leak secret sauce? Doubtful; it feels like its less about the specific attack of reverse-engineering Mythos, and more about being a general sophon against any model training at all; that Anthropic's official position is now that they're the only ones who should be training models.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#154

Earlier quoted context omitted.

> it won't just reject ML research, which I can understand I don't.

Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.

And they're hilariously pissy about it for a megacorp that did the same with the entire Internet and every library book they could get their hands on.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#155

I wonder how many millions they are wasting on putting up these guardrails when it's a completely useless exercise that is a speed bump at best.

If the guardrails were so useless, people wouldn't be complaining about them.

Murder is very (100%!) effective at preventing cancer. And yet, it is a useless method of preventing cancer.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#156

Earlier quoted context omitted.

It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…

That is for whatever it considers reverse-engineering the model to try to create a competing one.

No, it's not about reverse engineering. It targets ML research.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#157
I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably less verbose too, and the tok/s is unmatched for a wiki article pull.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#158
post #146

Earlier quoted context omitted.

Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.

Anthropic's claim was that Deepseek collected ~150k conversations. https://www.anthropic.com/news/detecting-and-preventing-dist... I think the extent of distillation by Deepseek specifically is overstated. For comparison, Minimax collected over 13m 'exchanges', which starts to sound a lot more like large-scale distillation.

Ah, dang it. My college professors warned me about this: the Wikipedia page I read the other day is wrong!

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#159
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

By saying they are 1 year ahead of their competition, it shows you don't know much about the pace LLM's and OpenAI's models.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#160

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

It royally pissed me off today by just continuing with credits without stopping to ask me if I was ok with it. Ran up $30 in extra charges while it was just flashing on the screen that it was doing that after I walked away to do something while it was humming along. It has always just told me I ran out of usage and had to wait before. Now? You’re just gonna pay extra because you left it unattended as you’ve done for…

You've already explicitly enabled extra usage in your account settings though, it is not on by default
Post reply on HN