Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

51–60 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#51

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

The problem is that Anthropic seems to be working up to the workflow one would naively want from AGI/some-god-like-entity.

The workflow would be; User asks for a thing. If it's a good thing, entity does the thing. If it's a naively bad idea, entity explains why you don't want that. If it's an actually evilly intended request, entity wags it's metaphorical finger or could even smite the user.

The problem is that flow isn't desirable if your entity isn't entirely god-like. It can bad even your entity is in ways rather far seeing.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#52
post #14

Earlier quoted context omitted.

If by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)

It was in the announcement, too. I’m 99% sure they edited it after they changed their mind, because I knew about it from reading that, and never opened the model card.

On the earliest web archive snapshot I can find [0], I do not see any mention of the safeguard/sabotage under discussion [1].

And to be clear, this isn't the safeguard where the model is explicitly downgraded to Opus, but rather where the Fable/Mythos model's "effectiveness" is transparently "limited" via "prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)".

[0]: https://web.archive.org/web/20260609173222/https://www.anthr...

[1]: https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-...

Re: Anthropic apologizes for invisible Claude Fable guardrails

#53
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Don't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll conveniently provide the framework for).

Dampened opinion on Anthropic is an understatement.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#54
post #10

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

What is "EA" in this context? I see a lot of people using this initialism.

"crypto bros" to a first approximation

Re: Anthropic apologizes for invisible Claude Fable guardrails

#55

Earlier quoted context omitted.

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Basically all critiques of Anthropic's policy moves on these topics boil down to people not believing the fundamental concerns are real, and often then going a step further to conclude that Anthropic doesn't actually believe their concerns either. If you believe Anthropic believes what they say they do, all of it makes sense.

What are you referring to? The cult belief that they are ushering in a machine god or that they strictly care about making as much money as humanely possibly while ignoring the absolutely destructive impacts these companies have had on society?

IMO they are using the cult messaging to distract the public so they take out all the oxygen in the room regarding people that care about the immediate impacts (climate exacerbation, ease of scamming, degrading job prospects, increasing income inequality).

Whenever real concerns are brought up against these companies they are always ignored while claiming the real concern is the fantasy of a machine god turning into skynet.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#56

Earlier quoted context omitted.

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#57

Earlier quoted context omitted.

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

Let's assume that Anthropic believes they're in an arms race to create a potentially dangerous technology, and they believe they're the best ones to win this race.

Unlike nuclear weapons, advancing in this arms race requires actually deploying the product over and over again. Deploying the product makes your advancements visible to your competitors.

It makes complete sense to try to limit the degree to which that's true.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#58

Earlier quoted context omitted.

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

Or if Google Chrome were blocking/degrading access to sites and services that might be useful to someone trying to make a competing web-browser.

P.S.: On reflection, it's even worse than that, because it'd trigger based on anything the user types or reads on any site. Someone mentions a "critical rendering path" and now you can't participate on that thread in the Blender forums.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#59

Earlier quoted context omitted.

I wonder who gets to decide which companies make important and critical software and which ones get the scraps later.

That would be Anthropic.

Well, Anthropic thinks it should be the Trump administration [1].

This whole business just keeps getting dumber.

1: https://darioamodei.com/post/policy-on-the-ai-exponential

Re: Anthropic apologizes for invisible Claude Fable guardrails

#60
Why would anyone defend Anthropic after this? Imagine falling for the DoW supply chain risk designation, and now this. This company is trying to ban powerful open models and restrict access to frontier models to slow everyone else down.

They just showed that they CAN do this right in front of you. Local open weight models are a necessity.

Post reply on HN