Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

291–300 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#291

Earlier quoted context omitted.

If a for profit company does a thing that could be motivated by profit or altruism, which of those 2 motivations do you think is most likely?

When they've repeatedly made decisions against their for profit nature, it changes the calculus a bit.

They haven't though. There's a long term plan here, and the goal is power and wealth. Short term moves that appear irrational turn out to be rational (from a greed perspective) when you factor in other considerations, like: Use their own AGI to create every software product on Earth and swallow the worlds economy. And we're kindly feeding their systems our codebases, IP and business decision-making so they can do exactly that.

Not a single thing Anthropic has done has been altruistic, and it never will be. It's all smoke and mirrors for the end goal.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#292

Earlier quoted context omitted.

The announcement elucidated this, and it's IMO worse than this. They don't downgrade to a cheaper model ([edit] for certain classes of offense they suspect you of). They sabotage the model's outputs in other, undisclosed, ways (specifically, "prompt modification, steering vectors, or parameter-efficient fine-tuning"). So, for example, they might load in a steering vector that just forgets the API to PyTorch. But it i…

It honestly explains so many issues I have been having, as I used it primarily for ML research (on my personal account, doing things not related to my job I should note). It would literally typo package names and spend huge amounts of time failing to setup simple environments…then do stupid things like set the learning rate to 1e-7, and use the eval set as training data.

just imagine if they made it sneaky. get things just subtly wrong enough that your training runs just never quite go as well as you think they should.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#294
post #176

Earlier quoted context omitted.

Consumer GPS is still disabled at high speeds. I would argue the analogy doesn't carry due to harm and error rate differences.

Yep a totally different use case and set of guardrails. There’s very little (not zero) consumer utility in GPS above say 15k feet AND 400 MPH or whatever the actual limit is. That’s basically tracking model rockets that are incidentally impacted and nothing else, from what I can think of.

It's also the sort of thing that has to have been thought up by someone with nothing better to do, given how ridiculous the premise is. You would have to assume the adversary is someone with the technology to build rockets, literally rocket science, but not the technology to build their own GPS receiver, which is simple 1970s radio technology?

Worse than that, it's 20th century radio technology in the 21st century when everyone has access to FPGAs and SDR.

The number of innocent people with model rockets or similar being negatively impacted by that rule is infinitely larger than the number of adversaries because the number of adversaries being impaired by it is zero.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#295
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#296
Maybe off-topic, but I'm also not happy about how they butchered my boy Opus 4.6. The model that could now hallucinates regularly.

Fable isn't even that great, not to mention it drinks token by the gallon for breakfast and keeps your data hostage for 30 days.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#297
post #285

Earlier quoted context omitted.

Or if GPU companies detected you were trying to train a model and injected intentional numerical errors.

Nvidia already did something similar with Lite Hash Rate (LHR), limiting performance on purpose just when running mining apps...

Well they did tell everyone explicitly and sell it as different SKUs. There's no Fable (Full ML) edition, just silent prompt injection.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#298

Earlier quoted context omitted.

Yep a totally different use case and set of guardrails. There’s very little (not zero) consumer utility in GPS above say 15k feet AND 400 MPH or whatever the actual limit is. That’s basically tracking model rockets that are incidentally impacted and nothing else, from what I can think of.

It's also the sort of thing that has to have been thought up by someone with nothing better to do, given how ridiculous the premise is. You would have to assume the adversary is someone with the technology to build rockets , literally rocket science, but not the technology to build their own GPS receiver, which is simple 1970s radio technology? Worse than that, it's 20th century radio technology in the 21st century w…

Errr I at least thought it would be easier to build a small, bad rocket than a precision GPS receiver. But I am not an expert.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#299
I don't want to be cynical, but I assume a third party we can trust has verified this model is actually this good?

I would think it would not be Anthropic, out of all the players, that is selling a lie hidden behind "I am sorry, I can't do that; it's too dangerous."

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#300
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

The mitigations against distillation are separate, and not what the OP is about at all.
Post reply on HN