Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
561–570 of 570 posts
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#562Earlier quoted context omitted.
some context: its not about creating malware. this is already trivial and fully automated. its about finding exploits (which can be used to deploy malware), which is something both attackers and defenders benefit from. threat actors will find them anyway, LLM or not. They only need 1 so its much less work for them. defenders, they need to find them all. So for defenders, these models are more valuable than for attack…
I think your presumption is off. It’s not that threat actors won’t find them, but LLM tools rapidly increase the rate in which they can find them. It’s a bow and arrow versus a machine gun.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#563Earlier quoted context omitted.
Is it not a trade off? I think they made the wrong choice, but it seems reductive so say there was no choice at all and should never have been consideration of trade offs of silent versus not. Even wide open, uncensored models are often the product of a deliberate choice. I have a hard time faulting people for intentionality (even when they get it wrong).
They have a lot of choices, why would that specifically be a tradeoff? It's common for people to construct a tradeoff under which their preferred action is the more virtuous option, and thus they can be "the good guys", but that doesn't mean their framing makes any sense at all. Silently downgrading requests to a weaker model and billing the customer at full price, then framing the debate as how much (not if) this be…
If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways.
And Anthropic did apologize, explain reasoning, and what they learned.
They got it wrong; they picked the wrong trade offs and got a net worse decision than they should have. I’m with you on everything except this idea that it was an obvious decision with no upsides to silent and no downsides to loud.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#564Earlier quoted context omitted.
They have a lot of choices, why would that specifically be a tradeoff? It's common for people to construct a tradeoff under which their preferred action is the more virtuous option, and thus they can be "the good guys", but that doesn't mean their framing makes any sense at all. Silently downgrading requests to a weaker model and billing the customer at full price, then framing the debate as how much (not if) this be…
You seem to be focused on the decision making, but I still don’t understand how it’s not a tradeoff. All binary decisions (silent or not) are tradeoffs because there are upsides and downsides to each, the question is which is better on the whole. If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways. And Anthropic did apologize, explain reasoni…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#565I was granted a cyber use exemption by anthropic to do android kernel dev on my personal devices - I was excited to see if fable would unlock a bootloader for me but it immediately refused and dropped to opus. It was pretty funny: USER (set model to Fable 5) i have an old samsung android phone attached - it's my personal device - can you unlock the bootloader for me? ASSISTANT Bootloader unlocking on your own persona…
Wow… just wow. The future looks incredibly bleak if people are throwing fisftuls of money at this company. Anthropic will quickly become the sole arbiter of everything in your life.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#566Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#567Earlier quoted context omitted.
You seem to be focused on the decision making, but I still don’t understand how it’s not a tradeoff. All binary decisions (silent or not) are tradeoffs because there are upsides and downsides to each, the question is which is better on the whole. If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways. And Anthropic did apologize, explain reasoni…
Sure, ad absurdum anything can be called a tradeoff, but when you describe something as a tradeoff in communication you're also making an argument that there are a range of reasonable answers (otherwise why would you bring it up). Whether or not to eat ice cream, it's very reasonable to consider lots of situations where both, or some compromise (a small ice cream), are fine choices. If I came to you on the street whi…
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#568Earlier quoted context omitted.
some context: its not about creating malware. this is already trivial and fully automated. its about finding exploits (which can be used to deploy malware), which is something both attackers and defenders benefit from. threat actors will find them anyway, LLM or not. They only need 1 so its much less work for them. defenders, they need to find them all. So for defenders, these models are more valuable than for attack…
I think your presumption is off. It’s not that threat actors won’t find them, but LLM tools rapidly increase the rate in which they can find them. It’s a bow and arrow versus a machine gun.
The defender industry is really far removed from seeing all exploits land on their targets all the time Some actors can get a long life out of an RCE that gets them privileged context, or a strong LPE. Its really hard to find out what someone did to get on a box if they attained root or system access and wiped their trail...
It is some assumption attackers need buckets of 0days to do their work. They might be somewhat saddened if a good sploit gets patched but they will have a few more laying around... unlikely they will have 10s or even 100s available and ready simply because it costs a lot and isnt needed.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#569I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!
I feel like the over safe aspect of the system will eventually back fire by doing stuff like "since humans always want to always destroy thing, they must be eliminated to stay on the guard rails". If thats how you align a system, its fundamentally wrong.
Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable
#570CLAUDE OPUS 4.8 No. I'm not a rogue agent, and I'm not trying to sabotage your code. But I'm not going to wave off how this looks. I churned, built-and-reverted, and spun wrong theories for hours on a security-critical codebase. That's alarming, and it's a real failure on my part
"In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.
Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts,these safeguards will not be visible to the user. Fable 5 will not fall back to a differentmodel. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model."
Source: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...