Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

421–430 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#422
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

The security guardrails are one thing but they extended it to AI work unrelated to security too to protect their lead.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#423
The underlying problem has not been resolved. People are required to trust Anthropic or anyone else. THAT is the big problem. I understand that some think this is a good trade-off; you may invest less time into writing code perhaps. But it is still a trade-off. I don't want to become dependent on Anthropic for anything.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#424

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> Fail cleanly.

Skynet does not fail.

It conquers.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#425
post #311
post #15

Earlier quoted context omitted.

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I think it's a big mistake to conflate the cyber (and bio) refusals with the LLM development refusals. I can sympathize with the argument for the cyber refusals - especially as a temporary measure - especially if Mythos is available to those trying to defend against vulnerabilities. The LLM development nerfing (and now refusals) is very different though. Anthropic has even said it isn't just for safety reasons: > Usi…

As we’ve seen with Fable, Mythos is more of a hype myth to justify the data retention and restrictions they added. Otherwise it’s just an incremental update of Opus. I can’t really say upgrade because the restrictions makes it a downgrade

Re: Anthropic apologizes for invisible Claude Fable guardrails

#430
Everyone with hostile intent runs local models.

Anyone with good intent, embracing the panopticon (of at least antroptics employees) works online. Thus the guardrails will always fail the protection goals by existing. They are purely for optics. The llm may as well make hostage negotiation smalltalk with you while you make secure software.

PS: To pay a cloud minimum-wage-employee for one "drop table weights" for mythos must be the equivalent of 5$ wrench to hit them over the head. https://imgs.xkcd.com/comics/security.png. Listen to that sound, that as if a whole ethics division got made redundant and unemployed.

Post reply on HN