I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
What is "EA" in this context? I see a lot of people using this initialism.
Anthropic apologizes for invisible Claude Fable guardrails
11–20 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#12I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
What is "EA" in this context? I see a lot of people using this initialism.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#13Re: Anthropic apologizes for invisible Claude Fable guardrails
#14*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails
(Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)
Re: Anthropic apologizes for invisible Claude Fable guardrails
#15I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#16Earlier quoted context omitted.
What is "EA" in this context? I see a lot of people using this initialism.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
Re: Anthropic apologizes for invisible Claude Fable guardrails
#17I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
This is the same exact industry that gives you paid usage limits as a unit-less percentage bar then gaslights customers every time the algorithm running that percentage bar changes or they lobotomize an existing model with increased quantization to squeeze a few more dollars out of existing hardware.
"Failing cleanly" might make their moated hype-machine look bad pre-IPO, so they certainly aren't going to do that voluntarily.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#18They didn't apologize for doing it, they are sorry they were caught doing it. They still nerf the model if your request is about AI development.
It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#19*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails
Re: Anthropic apologizes for invisible Claude Fable guardrails
#20But also, it isn’t the only huge mistake Anthropic has made in the last 48 hours. Having a sneaky data retention policy, while also giving companies no way to block Fable, is a massive problem. And it is ridiculous that Anthropic has so little respect for its customers. OpenAI should take advantage of this.