Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

11–20 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#11
post #10

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

What is "EA" in this context? I see a lot of people using this initialism.

Effective Altruism I think

Re: Anthropic apologizes for invisible Claude Fable guardrails

#12
post #10

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

What is "EA" in this context? I see a lot of people using this initialism.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

Re: Anthropic apologizes for invisible Claude Fable guardrails

#14

*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails

If by "got caught" you mean "published it in their system card paper".

(Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)

Re: Anthropic apologizes for invisible Claude Fable guardrails

#15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access.

Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#16
post #10

Earlier quoted context omitted.

What is "EA" in this context? I see a lot of people using this initialism.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#17

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Fail cleanly.

This is the same exact industry that gives you paid usage limits as a unit-less percentage bar then gaslights customers every time the algorithm running that percentage bar changes or they lobotomize an existing model with increased quantization to squeeze a few more dollars out of existing hardware.

"Failing cleanly" might make their moated hype-machine look bad pre-IPO, so they certainly aren't going to do that voluntarily.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#18

They didn't apologize for doing it, they are sorry they were caught doing it. They still nerf the model if your request is about AI development.

They didn't get "caught." It was published, by them, when they released Fable a few days ago. They were very clear about it.

It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#19

*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails

They didn’t get caught, they explicitly said they would do that in the announcement. I think it was both bad and a weird idea, but it certainly wasn’t sneaky.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#20
Invisible guardrails? Or purposeful sabotage if you use it for building AI capabilities?

But also, it isn’t the only huge mistake Anthropic has made in the last 48 hours. Having a sneaky data retention policy, while also giving companies no way to block Fable, is a massive problem. And it is ridiculous that Anthropic has so little respect for its customers. OpenAI should take advantage of this.

Post reply on HN