Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

251–260 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#251
post #83

Earlier quoted context omitted.

"Only I can save us". It's a classic tragedy and cautionary tale. The idea Anthropic was going to speed run AI so they could control the usage and make it "safe" for humanity was never altruistic; it was a HUGE FUCKING RED FLAG.

[flagged]

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#252

incredible marketing from anthropic with all the "it's too dangerous" bullshit

Agreed, it seems to be working and it's nonsense. I don't know why you're being downvoted.

"This information is too dangerous for you, so we'll just hold on to it.."

Thanks big brother, super anthropic of you!

The internet of '95 is looking back at us, with tears in its eyes.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#253
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Yeah I'd say this has been a big concern ever since it turned out immensely expensive training methods could create effective frontier models. So far at least, open source models have kept up better than I expected, but they definitely lag the top ones and there's no guarantee the gap doesn't widen further.

Imagine the software world if Linux never existed as an effective OS and Microsoft + Apple had completely controlled computer platforms for the past decades. I think it's almost certain that both companies would be even more profitable, and the tech industry would be vastly less free and more dysfunctional .

Re: Anthropic apologizes for invisible Claude Fable guardrails

#254
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

That level of control will be fleeting at best; as soon as the open models and competitors catch up they lose that influence

That's why Dario's advocating for making open weight models illegal and also saying we should stop the clock on model development amongst the large labs.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#256

Earlier quoted context omitted.

"Anthropic accused Chinese firms of 'industrial-scale distillation attacks' on its AI models." "Distillation involves training less capable models on more advanced ones’ output, and can be used illicitly to acquire powerful capabilities cheaply. The AI startup accused China’s DeepSeek, MiniMax, and Moonshot of generating 'over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts,'" https:…

It's definitely real, in the sense that it's a real violation of ToS. It could perhaps be used to guide a few narrow capabilities in very specific domains, given a model that's already most of the way there. But no, it's nowhere near the same as "stealing" a model outright, nor does it replace basic innovation in AI. And it's indistinguishable from practices that have long been common in the industry as a matter of f…

Oh, I agree distillation isn't stealing "outright" as in it's not theft of 100% of the model. But there's a reason they're doing it. I didn't say anything about Chinese labs innovating -- obviously they are.

What accounts for the difference between your attitude that distillation is no big deal, "common practice," yet Anthropic sees as it as a huge threat?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#257
post #3

Related. Others? Anthropic walks back policy that could have 'sabotaged' researchers using Claude - https://news.ycombinator.com/item?id=48485958 - June 2026 (30 comments) Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable - https://news.ycombinator.com/item?id=48478969 - June 2026 (488 comments) If Claude Fable stops helping you, you'll never know - https://news.ycombinator.com/item?id=…

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#258
post #3

Related. Others? Anthropic walks back policy that could have 'sabotaged' researchers using Claude - https://news.ycombinator.com/item?id=48485958 - June 2026 (30 comments) Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable - https://news.ycombinator.com/item?id=48478969 - June 2026 (488 comments) If Claude Fable stops helping you, you'll never know - https://news.ycombinator.com/item?id=…

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#259

Earlier quoted context omitted.

It's definitely real, in the sense that it's a real violation of ToS. It could perhaps be used to guide a few narrow capabilities in very specific domains, given a model that's already most of the way there. But no, it's nowhere near the same as "stealing" a model outright, nor does it replace basic innovation in AI. And it's indistinguishable from practices that have long been common in the industry as a matter of f…

Oh, I agree distillation isn't stealing "outright" as in it's not theft of 100% of the model. But there's a reason they're doing it. I didn't say anything about Chinese labs innovating -- obviously they are. What accounts for the difference between your attitude that distillation is no big deal, "common practice," yet Anthropic sees as it as a huge threat?

I never said that "it's no big deal". It's a clear-cut violation of ToS, and Anthropic are within their rights to care about that.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#260

Earlier quoted context omitted.

No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift. I participated in the internal bioweapons uplift test for Sonnet 3.7, and even then, one non-expert got huge uplift from the model [1]. I'd consider evals a lower bound of capabilities that can be elicited from a model. The team behind Biomni, a biomedical agen…

> No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift. Doesn't this simply amount to disagreeing about what counts as "meaningful" from a bio-safety POV? Also, even the ASL-3 deployment safeguards for Opus 4 and higher were always adopted as a mere matter of caution; it's not clear that even Anthropic believed at…

In normal bio, there are standardized biosafety levels, because without it there would be no standard agreement on what "meaningful" safety is. So yes, I do think there's ambiguity here.

But I don't think I've found any domain expert who thinks granting everyone raw access to the most capable models wouldn't meaningfully increase risk. OpenAI recently staffed a biological threat modeler to help quantify this risk.

(Edit: just saw your edit, this includes at Anthropic. ASL tiers were "rule-out" to exclude rather than "rule-in", so exact thresholds were murkier, but I think it's clear that models have passed that threshold by now.)

That said, there are clear steps and requirements to set up a BSL-2 or BSL-3 lab, and I think there should be similarly clear rules around model capabilties and access. The process for Anthropic and OpenAI is murky and still implictly gated on spend, which I think is holding back research.

For example, anyone who has access to a BSL-3 lab should have a clear and low-cost path to a model with corresponding capabilities, as long as they set up corresponding precautions for model access.

I think it would be a bad outcome for only frontier labs and a select few groups they choose to have access to the most capable models – which is sadly the precedent that's currently being set.

Post reply on HN