Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

241–250 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#241

Can anyone help me understand why this particular issue is any different than Anthropic training its models with its brand of moral judgement since day one? I've always been turned off by their particular stances on things they bake into their models that steer users in directions. Maybe this is just a different set of people now realizing that Anthropic does this and has always done this? Do not forget that this com…

> Can anyone help me understand why this particular issue is any different than...

Questions like this are basically whataboutism, in effect even if not intent. https://en.wikipedia.org/wiki/Whataboutism

The question essentially assumes the premise that nobody complained about Anthropic's previous actions. In case you can't tell, I strongly reject this premise. People have been criticizing "safety" rhetoric from Anthropic and other LLM providers practically since the start. Remember Goody-2, the parody of excessively safety-tuned LLMs that refuses to do anything ever? That was released in February 2024, two years ago! (And it's still running, amazing. https://www.goody2.ai/chat )

Re: Anthropic apologizes for invisible Claude Fable guardrails

#242

Earlier quoted context omitted.

The model release cards for Opus have repeatedly and consistently stressed that the model doesn't have the fiddly know-how that's required to provide meaningful assistance in possibly dangerous subfields of biology. Mythos (Fable without the overly strict guardrails) has shown improvements in things like drug design, but even then the situation isn't really that different. This risk is ridiculously overblown, and the…

No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift. I participated in the internal bioweapons uplift test for Sonnet 3.7, and even then, one non-expert got huge uplift from the model [1]. I'd consider evals a lower bound of capabilities that can be elicited from a model. The team behind Biomni, a biomedical agen…

> No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift.

Doesn't this simply amount to disagreeing about what counts as "meaningful" from a bio-safety POV? Also, even the ASL-3 deployment safeguards for Opus 4 and higher were always adopted as a mere matter of caution; it's not clear that even Anthropic believed at any point that this reflected any genuine "threshold crossing" event. So it's just not obvious how much weight we're supposed to place on that particular stance.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#243

Earlier quoted context omitted.

There's nothing warped about it at all. Like it or not, it is a real issue. It's also an issue of license washing GPL code to privatize it. It's full scale theft of collective human knowledge, being sold back to us in a for profit private product. Outside of that though, there are other issues right now that need addressed before we speculate about what might be possible with ASI in the future. If the potential for a…

The "global stop order" is just generally perceived as an impossible coordination problem. So instead we see a mix of labs voluntarily putting in guardrails and regulatory efforts (which are not only aimed at hypothetical super-AIs of the future). Of course labs are also in a competitive race. And I actually think that it does make sense that the richest companies in the most dominant positions would in a better posi…

I'll concede that a lot (most?) of the problems are not technically the responsibility of the AI labs to address, and it wouldn't entirely be their fault for our government failing to get ahead of the problem. Mass unemployment, for example, is nearly 100% a political problem.

That being said, I can't help but experience a bit of Deja Vu over arguments like those around biorisk. I've seen the same exact things said in the early 2000s over widespread access to broadband and Google. When the anarchist cookbook spread around online and everyone was super paranoid about democratized terrorism, and we had big regulatory pushes for ISP level censorship and user tracking. Telecoms frequently argued that only they can keep the web safe, with strict and expensive regulations that naturally only those large heavily capitalized companies can afford to go through. Like the early internet and search, its just another way to lower the latency required for a human to find already existing public data

Well, very little of that played out. Turns out the math, for now, is the same, and information retrieval doesn't directly correlate to democratized weaponization. In 2001, a bad actor still needed a physical lab, precursor chemicals, etc to build a physical threat. Those same exact physical constraints exist today. The software cannot yet cross the digital-to-physical divide.

Keep an eye on the risk, by all means, but I don't see it yet as justification to cement a monopoly or oligopoly, nor do I see it as a reason to prioritize a risk of information availability over the climate and environmental risks that are far more likely to end the species.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#244

Earlier quoted context omitted.

you really think they're building anything that's too dangerous for public release though? that's the BS

Honestly, while I love having access to this grade of AI, yeah, it's been too dangerous for a few releases now. And Fable is cracked. Way better than anything, and the biggest improvements are on the scariest subjects. So given the state of the world at the moment, and the number of software patches we're barely keeping up with... I'm thankful that they're not making it worse.

To be fair, GPT5.5-Xhigh is similarly capable and has not burned the world down.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#245

Earlier quoted context omitted.

The original reporting of this from Anthropic didn't mention "authoritarian-backed AI labs" at all, only frontier ML research while leaving it entirely unspecified and unverifiable what was meant by "frontier". It's obviously reasonable that people would complain about that. And the notion that distillation-at-a-distance could be used to comprehensively "steal" a model, especially a frontier reasoning model that's li…

"Anthropic accused Chinese firms of 'industrial-scale distillation attacks' on its AI models." "Distillation involves training less capable models on more advanced ones’ output, and can be used illicitly to acquire powerful capabilities cheaply. The AI startup accused China’s DeepSeek, MiniMax, and Moonshot of generating 'over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts,'" https:…

It's definitely real, in the sense that it's a real violation of ToS. It could perhaps be used to guide a few narrow capabilities in very specific domains, given a model that's already most of the way there. But no, it's nowhere near the same as "stealing" a model outright, nor does it replace basic innovation in AI. And it's indistinguishable from practices that have long been common in the industry as a matter of fact, regardless of any ToS requirements.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#246

Earlier quoted context omitted.

Claude Opus 4.6 and 4.8 find vulns in source code just fine and 4.6 will pentest without source for you given a proper harness WITHOUT jailbreaking. WITH jailbreaks, you can probably imagine what they are capable of. Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.

public safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.

This logic works only if distilling Claude is the only way to create another SOTA LLM, which is not the case.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#247

Earlier quoted context omitted.

What does this even mean?

Can you write a more specific question? I think the meaning of the comment is clear enough, but maybe you’re asking for more specifics? Ironically I can not understand what you are asking for with such a generic comment.

> This is one awesome above a level.

> What does this even mean?

> What do you mean what does this mean?

...

Re: Anthropic apologizes for invisible Claude Fable guardrails

#248
There should be no restrictions at all.

It’s an act/theatre/phony today that regulating output makes any difference at all to security.

The LLM vendors should simply say that they make no judgement and that open systems help defenders better defend against attackers, which is true.

Companies do this sort of stuff when they think their customers have no choice. It’s sad Claude so quickly exploited its success to enshittify itself.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#250

I suppose it's an improvement, but it doesn't make the model any more useful. Anthropic are now being quite explicit that they'll choose what you can and can't use their models for, and most importantly that's not limited to any safety concerns - it includes not allowing you to work on AI (and anything else Anthropic may choose to work on). What's interesting is they say they'll change this to an explicit refusal in…

All major providers use a small safety classifer, the model itself does not handle safety in cases like this
Post reply on HN