Live data from Hacker News

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

techcrunch.com

351–360 of 570 posts

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#351
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

I work on open source text-to-image finetuning of open source models like zimage/flux2 klein 4b and inference time latency optimization. The moment I read the silent treatment, I went ahead and cancelled my subscription too since I would never know whether the models they launch will silently corrupt my output. This is totally unacceptable. There is a big difference between silent / flagged if you are doing ml research but not at frontier capability.

This goes on to show that - All that interpretability / safety research they are doing can also be weaponized against customers (steering vectors, intent classification, ...) in the name of safety from malicious actors. - If they deem profitable, they might nerf to original model and its training data for ml research at a bulk scale and then they won't even have to announce it so long as the overall benchmark score stays high enough.

As the IPOs get closer, they can do whatever they want to assure the investors that they have a moat that can not be crossed over by their own products. Considering this affects all ML researchers/students at universities, smaller scale research labs, this is just "cutting the branch you are sitting on".

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#352
post #96
post #56

Earlier quoted context omitted.

I assume they’re encrypted/DRM’ed when deployed on inference hardware, so only core researchers/sec admins would potentially have some access to unprotected weights, and they are far too well paid to risk it leaking the model

I wouldn't be surprised if they encrypt them at rest, but at some point the weights have to be loaded into vram.

Newer NVidia cards (H100 and up) support both in-memory model encryption and ‘trusted’ execution environment/remote attestation, not sure how widely used in frontier model deployments, but at least vendor claimed perf overhead is ‘3%’ [0]

[0] https://www.spheron.network/blog/confidential-gpu-computing-...

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#353

Earlier quoted context omitted.

> speed bumps work Yea. To slow you down. They don't prevent you from getting somewhere.

> To slow you down. They don't prevent you from getting somewhere Again, yeah. That's how fences work, too. And alarm systems. Pretty much anything that isn't foolproof. Pointing out that a defence is surmountable isn't a rejection of it per se .

Fences and speed bumps are hilarious defences if we are supposed to believe AI companies about the dangers of this technology.

Having no safeguards is probably safer than having safeguards which do nothing but create a false sense of security.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#354
post #146

Earlier quoted context omitted.

Anthropic has already been burned before on this. DeepSeek was trained on million of conversations with Claude. And DeepSeek created thousands of free accounts to burn all this compute at their expense.

Anthropic's claim was that Deepseek collected ~150k conversations. https://www.anthropic.com/news/detecting-and-preventing-dist... I think the extent of distillation by Deepseek specifically is overstated. For comparison, Minimax collected over 13m 'exchanges', which starts to sound a lot more like large-scale distillation.

If that's all it took to make Deepseek so good, I'll gladly ship High-Flyer all my personal 150k claude/chatgpt conversations in exchange for Deepseek 5 (and a rack of B200s or Ascend chips)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#355
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

This is different to the cyber limitations though.

To be precise - it makes the "won't work on frontier machine learning" refusal the same as the "won't work on cyber security" refusal (instead of the way it previously would work on frontier machine learning problems but give sub-optimal answers without informing the user)

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#356
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

Non-paywalled: https://archive.md/yxYhU

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#357
post #95

Earlier quoted context omitted.

Why is this surprising or a problem?! It's a model demo, & their reasoning is reasonable and fair. Why all this drama.

Some people find Anthropic's special blend of paternalism and random incompetence tiresome.

"I will push back and say" it's only paternalism if it's about helping the user's not harm themselves.

This is about societal impacts, not wanting their models to be used by some people against other people, as a weapon.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#358
post #350

Earlier quoted context omitted.

Because most people in tech never took a philosophy course or an ethics course and think that tech is obviously a good for the world and that there are no downsides to advancing tech. So any efforts that try to apply ethics to it are overreaching, ignorant, and futile in the face of the good that is tech!

So i have big news for you my friend as i'm not sure you understand such courses. Taking an ethics course won't make you a more ethical person.. and taking a philosophy course neither.

You're being too literal, they're saying people are not thinking with a philosophically interested mind, which is blatantly the case here, their point stands.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#359

Earlier quoted context omitted.

Tech demo + theres the ability to provide feedback right at the answer interface if using the UI. Provide feedback in the negative, a brief explanation, and move on with your day. It will improve with feedback, not with whinging into the void.

Ironically making a stink about it online is likely to have a larger impact then using their dedicated feedback or support channels (which go to claude, not a person)

the feedback is for something mindless though, "we don't care about societal harms". I wonder the overlap between these commenters and tech maga people, eg crypto bros & Elon stans.

Re: Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

#360
post #284

News just broke in this Wired story: "Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude" https://www.wired.com/story/anthropic-responds-to-backlash-o... > “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” Sounds like the wides…

To late. I canceled my Max subscription. The idea they would even do this is so destroyed any remaining trust. Why would I pay them 1000s of dollars in extra usage per month for something they could still be doing behind the scenes? Any errors previously chalked up to thinking effort or other backend changes? Maybe it was intentional prompt injection the entire time.

OpenAI has a real opportunity to do some sort of "we don't maliciously alter your prompt and nerf the model" with some form of verification, when they release the next model.

But if Anthropic gets their way with regulatory capture, this could be the only future we'll see.

To think that they didn't expect the backlash speaks volumes about how much shady things they're doing which is not publicly known.

Post reply on HN