Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

561–570 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#562

Earlier quoted context omitted.

some context: its not about creating malware. this is already trivial and fully automated. its about finding exploits (which can be used to deploy malware), which is something both attackers and defenders benefit from. threat actors will find them anyway, LLM or not. They only need 1 so its much less work for them. defenders, they need to find them all. So for defenders, these models are more valuable than for attack…

I think your presumption is off. It’s not that threat actors won’t find them, but LLM tools rapidly increase the rate in which they can find them. It’s a bow and arrow versus a machine gun.

They can also potentially allow said issues to be found and fixed more quickly - and also allow teams to implement deeper security boundaries throughout their systems such that one big steel door getting compromised does not lead to everything being easily available.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#563

Earlier quoted context omitted.

Is it not a trade off? I think they made the wrong choice, but it seems reductive so say there was no choice at all and should never have been consideration of trade offs of silent versus not. Even wide open, uncensored models are often the product of a deliberate choice. I have a hard time faulting people for intentionality (even when they get it wrong).

They have a lot of choices, why would that specifically be a tradeoff? It's common for people to construct a tradeoff under which their preferred action is the more virtuous option, and thus they can be "the good guys", but that doesn't mean their framing makes any sense at all. Silently downgrading requests to a weaker model and billing the customer at full price, then framing the debate as how much (not if) this be…

You seem to be focused on the decision making, but I still don’t understand how it’s not a tradeoff. All binary decisions (silent or not) are tradeoffs because there are upsides and downsides to each, the question is which is better on the whole.

If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways.

And Anthropic did apologize, explain reasoning, and what they learned.

They got it wrong; they picked the wrong trade offs and got a net worse decision than they should have. I’m with you on everything except this idea that it was an obvious decision with no upsides to silent and no downsides to loud.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#564

Earlier quoted context omitted.

They have a lot of choices, why would that specifically be a tradeoff? It's common for people to construct a tradeoff under which their preferred action is the more virtuous option, and thus they can be "the good guys", but that doesn't mean their framing makes any sense at all. Silently downgrading requests to a weaker model and billing the customer at full price, then framing the debate as how much (not if) this be…

You seem to be focused on the decision making, but I still don’t understand how it’s not a tradeoff. All binary decisions (silent or not) are tradeoffs because there are upsides and downsides to each, the question is which is better on the whole. If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways. And Anthropic did apologize, explain reasoni…

Sure, ad absurdum anything can be called a tradeoff, but when you describe something as a tradeoff in communication you're also making an argument that there are a range of reasonable answers (otherwise why would you bring it up). Whether or not to eat ice cream, it's very reasonable to consider lots of situations where both, or some compromise (a small ice cream), are fine choices. If I came to you on the street while you were holding your ice cream, told you I was going to take your ice cream and eat it, and after your protest changed my mind and said that was the "wrong balance" in how much I took, you could very fairly conclude that I was the sort of person who thinks it's ok to steal people's ice cream (maybe just not today). Now maybe there was some great miscommunication, and I thought you were done with it, but then I wouldn't say I made the wrong tradeoff, I would just tell you I misunderstood the situation and hope you'd invite me to ice cream another time. Anthropic in this case is saying they think there's a balance in how much of the service customers pay for that they should deliver, as clarified in their Wired statement. That's a choice.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#565

I was granted a cyber use exemption by anthropic to do android kernel dev on my personal devices - I was excited to see if fable would unlock a bootloader for me but it immediately refused and dropped to opus. It was pretty funny: USER (set model to Fable 5) i have an old samsung android phone attached - it's my personal device - can you unlock the bootloader for me? ASSISTANT Bootloader unlocking on your own persona…

Wow… just wow. The future looks incredibly bleak if people are throwing fisftuls of money at this company. Anthropic will quickly become the sole arbiter of everything in your life.

thankfully we will have a competing frontier model from the next lab shortly I'm sure and I hope they are reading this feedback and swing theirs the other direction.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#567

Earlier quoted context omitted.

You seem to be focused on the decision making, but I still don’t understand how it’s not a tradeoff. All binary decisions (silent or not) are tradeoffs because there are upsides and downsides to each, the question is which is better on the whole. If I’m deciding whether or not to eat ice cream, there are trade offs involved because I can’t simultaneously have it both ways. And Anthropic did apologize, explain reasoni…

Sure, ad absurdum anything can be called a tradeoff, but when you describe something as a tradeoff in communication you're also making an argument that there are a range of reasonable answers (otherwise why would you bring it up). Whether or not to eat ice cream, it's very reasonable to consider lots of situations where both, or some compromise (a small ice cream), are fine choices. If I came to you on the street whi…

(they also reset usage for many accounts after the Mythos/Fable rollback, which is great)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#568

Earlier quoted context omitted.

some context: its not about creating malware. this is already trivial and fully automated. its about finding exploits (which can be used to deploy malware), which is something both attackers and defenders benefit from. threat actors will find them anyway, LLM or not. They only need 1 so its much less work for them. defenders, they need to find them all. So for defenders, these models are more valuable than for attack…

I think your presumption is off. It’s not that threat actors won’t find them, but LLM tools rapidly increase the rate in which they can find them. It’s a bow and arrow versus a machine gun.

i dont think so perse simply because attackers dont need a lot of the exploits to be 'fired' continually at targets. They need few reliable and unknown ones.

The defender industry is really far removed from seeing all exploits land on their targets all the time Some actors can get a long life out of an RCE that gets them privileged context, or a strong LPE. Its really hard to find out what someone did to get on a box if they attained root or system access and wiped their trail...

It is some assumption attackers need buckets of 0days to do their work. They might be somewhat saddened if a good sploit gets patched but they will have a few more laying around... unlikely they will have 10s or even 100s available and ready simply because it costs a lot and isnt needed.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#569
post #170

I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!

I feel like the over safe aspect of the system will eventually back fire by doing stuff like "since humans always want to always destroy thing, they must be eliminated to stay on the guard rails". If thats how you align a system, its fundamentally wrong.

Achievement Unlocked: Invent Skynet

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#570
Today, working on remote attestation for a distributed system...

CLAUDE OPUS 4.8 No. I'm not a rogue agent, and I'm not trying to sabotage your code. But I'm not going to wave off how this looks. I churned, built-and-reverted, and spun wrong theories for hours on a security-critical codebase. That's alarming, and it's a real failure on my part

"In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts,these safeguards will not be visible to the user. Fable 5 will not fall back to a differentmodel. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. When these interventions are active, we expect them to have minimal behavioral impact on the model except to limit its effectiveness in developing frontier LLMs. Claude will still respond helpfully to user requests. We’ll continue to improve the precision of our detection methods following the launch of this model."

Source: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3...

Post reply on HN