Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

21–30 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#21
I'm surprised they didn't do this the first time around. Like, a user says they forgot their password and you tell them they don't actually have an account, that's an information disclosure vulnerability. Not automatically falling back to Opus just lets the "attacker" know they are bumping against the guardrails and they need to try a different strategy.

It's Anthropic's product and they can do what they want, but my concern is what happens if Fable's product team decides that they can route 25% of traffic to Opus, bill it as Fable, and max their KPIs. That just doesn't sit right.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#22

They didn't apologize for doing it, they are sorry they were caught doing it. They still nerf the model if your request is about AI development.

They didn't get "caught." It was published, by them, when they released Fable a few days ago. They were very clear about it. It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.

No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior.

https://xcancel.com/hammer_mt/status/2064839924398825798

Re: Anthropic apologizes for invisible Claude Fable guardrails

#23
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

Claude Opus 4.6 and 4.8 find vulns in source code just fine and 4.6 will pentest without source for you given a proper harness WITHOUT jailbreaking. WITH jailbreaks, you can probably imagine what they are capable of.

Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#24
post #14

*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails

If by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)

It was in the announcement, too. I’m 99% sure they edited it after they changed their mind, because I knew about it from reading that, and never opened the model card.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#26
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

exactly for cybersecurity the failure was visible. It was not visible for "Frontier" ML Research. The argument of headstart in it security is no feasible here.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#27
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I wonder who gets to decide which companies make important and critical software and which ones get the scraps later.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#28

Earlier quoted context omitted.

They didn't get "caught." It was published, by them, when they released Fable a few days ago. They were very clear about it. It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.

No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798

Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#30
post #21

I'm surprised they didn't do this the first time around. Like, a user says they forgot their password and you tell them they don't actually have an account, that's an information disclosure vulnerability. Not automatically falling back to Opus just lets the "attacker" know they are bumping against the guardrails and they need to try a different strategy. It's Anthropic's product and they can do what they want, but my…

It failed visible for it security and bio/chemistry stuff. It sabotaged invisible for "frontier" ML research. Its not a switch to a cheaper model. They tried to actively harm progress.
Post reply on HN