Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

161–170 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#161
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

One year ahead of it's competition in what exactly? Vibe coding? From Opus 4.7 onwards each following model is becoming less useful as an assistant and turning you as the assistant. But I guess that's normal when it's trained to pass benchmarks end to end. In fact it has become extremely good at pushing against feedback with extremely convincing and intelligent takes, even when it's completely wrong. I have extensive…

They def not 1 year ahead, at most 2 weeks ahead until Openai releases theirs. This guy def a Anthropic shill and probably doesn't use any other LLMs.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#162

I wear a few hats, but as a chemist and I'm not happy with fable. As a statistician I'm not happy with fable. As a data scientist I am not happy with fable. As an academic and a researcher I am not happy with fable. It's useless. I'd be surprised if anyone can get any output from it that couldn't easily be replaced with a search from wikipedia. Given how verbose claude models have become, wiki articles are probably l…

I’ve been working on a rather complex mapping project and have been getting MUCH better results with Fable than Opus.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#163

I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!

I wonder if it blurred the image or something before passing it to Opus...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#165
post #129

Earlier quoted context omitted.

The thing that I keep thinking about is the accounting / charging when it downgrades automatically. Do they adjust the price of the api request so that only the tokens that were utilized by fable get charged at that price and the remaining tokens that the cheaper / nerfed (fable) model utilizes get charged at that price? If the answer is no, could that be construed as fraud?

Their goal is to downgrade people who are violating their TOS, so I think they'd have some argument there. I have no idea how they'll deal with inevitable false positives, especially given how oversensitive most of the other triggers are.

The challenge is the examples they’ve mentioned (distributed training infra? ML acceleration techniques?) go beyond what’s prohibited by their ToS and is like a catch net.

I would wager the majority of ML and data science work in the world aren’t frontier LLM development.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#166
post #5

The strangest part is that it won't just reject ML research, which I can understand, it will sabotage it silently by using a worse model without revealing it is doing so. It's just an insane level of deception and trust destruction for a company that at most is like 1 year ahead of its competition. Edit; to be clear they tell you when they degrade it for cybersecurity and bio

Can you imagine if AMD or Intel throttled your cpu if it detected you were working on "cybersecurity" or if you were designing a cpu?

It would suck, but guardrails on new technologies like this aren't unheard of. It's like when consumer GPS used to stop working at very high speeds because they didn't want people to use it for missile guidance systems.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#167
post #138

At least Anthropic weren't lying when they said only a week ago or so "No one has figured out guardrails yet", because they apparently haven't either and Fable simply flat out rejects anything remotely connected to biology or security, no matter how trivial.

> At least Anthropic weren't lying when they said only a week ago or so "No one has figured out guardrails yet"

Anthropics guardrails are the TSA saying "take off your shoes" while failing every test. https://oversightdemocrats.house.gov/news/press-releases/new...

Anthropic owns the TOS... "If we think your involved in criminal activity were turning all your history over to the FBI/CIA/NSA/Local police". Then if their tooling was so good offering the same agency analysis tools to aid their experts in making some sort of decision.

But their detection isnt that good, and their analysis isnt either... this is pure theater, to create buzz (no such thing as bad press) and make their tool look far better than it is.

The reality is that, they arent even looking for the vectors that pose some of the largest risks in the modern era. And when someone uses it to do something terrible, they did not think of they are going to look dumb.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#168

Earlier quoted context omitted.

It's from the model card: > unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). https://www-cdn.anthropic.com/d00db56fa754a1…

That is for whatever it considers reverse-engineering the model to try to create a competing one.

No, that’s for “frontier LLM development” which somehow includes examples like distributed training infra.

Based on how sensitive the classifers are, any data scientist / MLE is probably going to encounter cases where some silent degradation happens and you never know about it.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#169
post #121

These guardrails are solely a reason for using your data for training purposes. Every flagged message can be used for training.

> We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose Whatever problem we might have with them, they explicitly say that they do not do this in the launch post.

"We won’t use this data to train new Claude models"

What about non-Claude models?

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#170

I tried asking Fable 5 to identify the fungus in a picture I uploaded of one of my wife's plants. Apparently it thought I was trying to build a bioweapon. Opus answered it (yellow dog vomit fungus). Now I can spread the spores and take over the world!

I feel like the over safe aspect of the system will eventually back fire by doing stuff like "since humans always want to always destroy thing, they must be eliminated to stay on the guard rails". If thats how you align a system, its fundamentally wrong.
Post reply on HN