Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

261–270 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#261
post #95

Earlier quoted context omitted.

[flagged]

What backward logic is this? PRC doesn't give a fuck about how US regulates AI companies. Pushing more regulation would ensure that Chinese companies catch up sooner. If you think otherwise you need to think harder.

It's a good thing you weren't in charge of nuclear arsenals during the Cold War, sounds like your approach would have been unchecked proliferation.

Fortunately developing frontier models takes immense amounts of specific resources and knowledge. There are only a handful of companies capable of developing new cutting edge models. This is an area a few governments absolutely could coordinate on and regulate, if they were so inclined.

Obviously the current US administration is completely lacking both the will and competence to actually negotiate an agreement like that with China, and who knows if Xi would even be interested. But with different leadership we actually could be reducing our existential risks in this area much more than we are. Just like having a few thousands nukes across several countries isn't totally safe, but it's a heck of a lot safer than hundreds of thousands of nukes spread across a hundred countries.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#262

Earlier quoted context omitted.

> No, Anthropic's model cards have claimed that the models don't show considerably more uplift than previous ASL-3 models, which already showed material uplift. Doesn't this simply amount to disagreeing about what counts as "meaningful" from a bio-safety POV? Also, even the ASL-3 deployment safeguards for Opus 4 and higher were always adopted as a mere matter of caution; it's not clear that even Anthropic believed at…

In normal bio, there are standardized biosafety levels, because without it there would be no standard agreement on what "meaningful" safety is. So yes, I do think there's ambiguity here. But I don't think I've found any domain expert who thinks granting everyone raw access to the most capable models wouldn't meaningfully increase risk. OpenAI recently staffed a biological threat modeler to help quantify this risk. (E…

> But I don't think I've found anyone who is a domain expert who thinks granting everyone access to raw modes wouldn't meaningfully increase risk.

It depends how capable these raw models are. Biology as a field depends most on real-world knowledge, which is an expensive capability for open models targeting widespread deployment. It's quite plausible that even Opus 4 would be a lot more capable in these domains than the best universally accessible "raw models" today, quite unlike other domains such as coding or pure math. The securebio.org benchmark has spotty representation of openly available models, but it does show Kimi 2.5 being no more capable than GPT 5 mini, and clearly below o4-mini and Opus 4.0; which may be a plausible summary of where things stand today.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#263
post #95

Earlier quoted context omitted.

[flagged]

Right now the PRC is looking like the adult in the room. They also have a view of how AI should work that's smaller and more worker centric rather than trying to create superintelligent worker replacements. The PRC (like any superpower) has done some bad shit, but if you're going to paint them as the bad guy keep in mind the USA has a long, long history of genocide, slavery, overthrowing foreign governments for corpo…

> Right now the PRC is looking like the adult in the room

Only if you ignore history.

Didn't the PRC violate every known labor/enviromnetal/human-rights standard to become the top in manufacturing?

https://matthewekahn.substack.com/p/what-role-did-regulation...

Re: Anthropic apologizes for invisible Claude Fable guardrails

#264

Earlier quoted context omitted.

Right now the PRC is looking like the adult in the room. They also have a view of how AI should work that's smaller and more worker centric rather than trying to create superintelligent worker replacements. The PRC (like any superpower) has done some bad shit, but if you're going to paint them as the bad guy keep in mind the USA has a long, long history of genocide, slavery, overthrowing foreign governments for corpo…

> Right now the PRC is looking like the adult in the room Only if you ignore history. Didn't the PRC violate every known labor/enviromnetal/human-rights standard to become the top in manufacturing? https://matthewekahn.substack.com/p/what-role-did-regulation...

The US did the same thing. Environmentalist and workers rights movements date back to the 19th century. China's position on this is that the western nations that already developed are trying to pull the ladder they used up and wag a finger with false morality with the intent of maintaining global hegemony.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#265
post #94
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Google has been doing the same thing for longer than Anthropic[0]. To protect their models from distillation attacks, they silently will downgrade the model's performance to essentially poison your training data without your knowledge. A bit different than Anthropic refusing to assist with any AI development at all, but it's in the same vein and seems not widely known. edit: reading the whole series of Google's AI Th…

Thanks for flagging this. This is interesting

Re: Anthropic apologizes for invisible Claude Fable guardrails

#266

Earlier quoted context omitted.

In normal bio, there are standardized biosafety levels, because without it there would be no standard agreement on what "meaningful" safety is. So yes, I do think there's ambiguity here. But I don't think I've found any domain expert who thinks granting everyone raw access to the most capable models wouldn't meaningfully increase risk. OpenAI recently staffed a biological threat modeler to help quantify this risk. (E…

> But I don't think I've found anyone who is a domain expert who thinks granting everyone access to raw modes wouldn't meaningfully increase risk. It depends how capable these raw models are. Biology as a field depends most on real-world knowledge, which is an expensive capability for open models targeting widespread deployment. It's quite plausible that even Opus 4 would be a lot more capable in these domains than t…

That's a good clarification. I've updated my comment to the "most capable models" to refer to the most recent releases.

And sure, and I love open models – I spent much of the past couple months doing additional RL on Qwen 3.6 35B A3B, Gemma 4, Kimi K2.6, and GLM 5.1. Without these open models, I'd be forced to do my research inside a frontier lab.

There's a balance to strike here, but I don't think the biological risk is overplayed. It would be very easy to accidentally cross the threshold of "meaningful" without adequate safeguards, and then be unable to undo what you've released to the world.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#267

Earlier quoted context omitted.

[flagged]

Safety from what? Competitors? That sounds like a product decision. They're puking on any requests that could be used to create LLMs or competitive products.

Anything to prevent mecha ai hitler. At all costs

Re: Anthropic apologizes for invisible Claude Fable guardrails

#268
post #261

Earlier quoted context omitted.

What backward logic is this? PRC doesn't give a fuck about how US regulates AI companies. Pushing more regulation would ensure that Chinese companies catch up sooner. If you think otherwise you need to think harder.

It's a good thing you weren't in charge of nuclear arsenals during the Cold War, sounds like your approach would have been unchecked proliferation. Fortunately developing frontier models takes immense amounts of specific resources and knowledge. There are only a handful of companies capable of developing new cutting edge models. This is an area a few governments absolutely could coordinate on and regulate, if they we…

> It's a good thing you weren't in charge of nuclear arsenals during the Cold War

You know how many nukes Soviet had right at its peak? Hint: much more than the US by the time. Non proliferation didn't stop Soviet from building more nukes at all. And it's not going to stop China from pouring more computing power into AI. History is a really good lesson.

The whole point of non-proliferation is to ensure that big boys like the US and Soviet can bully smaller guys like Venezuela and Ukraine. In this regard, non-proliferation is the most successful foreign policy ever. But it didn't win the cold war and a similar policy over AI will definitely not win the AI race (if it's a race worth winning is another issue.)

Re: Anthropic apologizes for invisible Claude Fable guardrails

#269
post #194

Earlier quoted context omitted.

The first half of your answer presupposes some platonic utilitarian calculus that, if it were applied correctly, would yield moral outcomes. This is very hard to believe. If I look at notable/well-known examples of EA-affiliated people, it is hard to skip by members such as SBF. Did he correctly apply the utilitarian calculus? It is relatively easy to take the proceeds of a massive fraud, buy a relatively small (as a…

Yeah, I didn't mean to downplay how hard it is to apply the utilitarian calculus or even to suppose that the bare doctrine of utilitarianism resolves questions about what the ultimate good we should be trying to maximize is. I basically agree that utilitarianism is not a complete recipe for how to live. I just think that it probably gives the correct answer in cases where we can see clearly how to apply it because I'…

This kind of reasoning leads you to reasoning that if he was an ineffective fraudster, it would be less moral, as he would have bought less mosquito nets. So it’s not only moral to do fraud, but you most extremely competently do fraud.

I think this being a reasonable utilitarian point to make is not a point in utilitarianism’s favor.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#270

Earlier quoted context omitted.

> Right now the PRC is looking like the adult in the room Only if you ignore history. Didn't the PRC violate every known labor/enviromnetal/human-rights standard to become the top in manufacturing? https://matthewekahn.substack.com/p/what-role-did-regulation...

The US did the same thing. Environmentalist and workers rights movements date back to the 19th century. China's position on this is that the western nations that already developed are trying to pull the ladder they used up and wag a finger with false morality with the intent of maintaining global hegemony.

> The US did the same thing.

Except that there were no global standards at the time. You can't point to any single country and say they were doing worse. They all were bad.

But China actively flouted established international norms. Now that is behind in AI it is clamoring for controls for others.

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3692695

> are trying to pull the ladder they used up

Every country spies and steals but it is the scale we are talking about. China does it at a scale that dwarfs any historical or current comparisons.

China doesn't have any grounds here when they turn around and complain about India copying its playbook:

https://economictimes.indiatimes.com/industry/renewables/chi...

Post reply on HN