Earlier quoted context omitted.
I did in my comment above.
You said these groups have access to LLMs. So what? Mythos/Fable are a step change above most LLMs. Responsibly limiting access and easing it up over time safely is the sane move.
Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
151–160 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#152Earlier quoted context omitted.
> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.
From the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#153Earlier quoted context omitted.
It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…
That is for whatever it considers reverse-engineering the model to try to create a competing one.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#154Earlier quoted context omitted.
> it won't just reject ML research, which I can understand I don't.
Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#155I wonder how many millions they are wasting on putting up these guardrails when it's a completely useless exercise that is a speed bump at best.
If the guardrails were so useless, people wouldn't be complaining about them.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#156Earlier quoted context omitted.
It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…
That is for whatever it considers reverse-engineering the model to try to create a competing one.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#157Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#158Earlier quoted context omitted.
Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.
Anthropic's claim was that Deepseek collected ~150k conversations. https://www.anthropic.com/news/detecting-and-preventing-dist... I think the extent of distillation by Deepseek specifically is overstated. For comparison, Minimax collected over 13m 'exchanges', which starts to sound a lot more like large-scale distillation.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#159The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#160Earlier quoted context omitted.
The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?
It royally pissed me off today by just continuing with credits without stopping to ask me if I was ok with it. Ran up $30 in extra charges while it was just flashing on the screen that it was doing that after I walked away to do something while it was humming along. It has always just told me I ran out of usage and had to wait before. Now? You’re just gonna pay extra because you left it unattended as you’ve done for…