Earlier quoted context omitted.
What is "EA" in this context? I see a lot of people using this initialism.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
Anthropic apologizes for invisible Claude Fable guardrails
31–40 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#32Earlier quoted context omitted.
What is "EA" in this context? I see a lot of people using this initialism.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
Re: Anthropic apologizes for invisible Claude Fable guardrails
#33I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#34Earlier quoted context omitted.
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
I wonder who gets to decide which companies make important and critical software and which ones get the scraps later.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#35I'm surprised they didn't do this the first time around. Like, a user says they forgot their password and you tell them they don't actually have an account, that's an information disclosure vulnerability. Not automatically falling back to Opus just lets the "attacker" know they are bumping against the guardrails and they need to try a different strategy. It's Anthropic's product and they can do what they want, but my…
It failed visible for it security and bio/chemistry stuff. It sabotaged invisible for "frontier" ML research. Its not a switch to a cheaper model. They tried to actively harm progress.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#36Re: Anthropic apologizes for invisible Claude Fable guardrails
#37*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails
If by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)
They could have simply told people "we do not permit using Claude models to perform frontier AI research," which is defensible from a policy point of view. This particular usage of their products requires no deception, nor hiding information prevent abuse.
However, instead, they chose for some reason to publicly display a morally poor way to execute a reasonable business decision (preventing abuse, defending your business interests, etc.)
Re: Anthropic apologizes for invisible Claude Fable guardrails
#38God bless the Chinese companies releasing true open source models. Imagine a world without them, we would be at the mercy of unscrupulous people.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#39Earlier quoted context omitted.
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
I wonder who gets to decide which companies make important and critical software and which ones get the scraps later.
The answer is, the organization making the powerful tool. The people in charge of Anthropic.
Not only that, but they've also written at length about exactly what their opinions and values are: https://darioamodei.com/
You may not agree with the decisions that they make, but they're hardly mysterious. Not something to wonder about.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#40Earlier quoted context omitted.
No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798
Making excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.