Anthropic apologizes for invisible Claude Fable guardrails
421–430 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#422I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#423Re: Anthropic apologizes for invisible Claude Fable guardrails
#424I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…
Skynet does not fail.
It conquers.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#425Earlier quoted context omitted.
I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
I think it's a big mistake to conflate the cyber (and bio) refusals with the LLM development refusals. I can sympathize with the argument for the cyber refusals - especially as a temporary measure - especially if Mythos is available to those trying to defend against vulnerabilities. The LLM development nerfing (and now refusals) is very different though. Anthropic has even said it isn't just for safety reasons: > Usi…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#426Re: Anthropic apologizes for invisible Claude Fable guardrails
#427Re: Anthropic apologizes for invisible Claude Fable guardrails
#428Re: Anthropic apologizes for invisible Claude Fable guardrails
#429Re: Anthropic apologizes for invisible Claude Fable guardrails
#430Anyone with good intent, embracing the panopticon (of at least antroptics employees) works online. Thus the guardrails will always fail the protection goals by existing. They are purely for optics. The llm may as well make hostage negotiation smalltalk with you while you make secure software.
PS: To pay a cloud minimum-wage-employee for one "drop table weights" for mythos must be the equivalent of 5$ wrench to hit them over the head. https://imgs.xkcd.com/comics/security.png. Listen to that sound, that as if a whole ethics division got made redundant and unemployed.