Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

401–410 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#402
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Same. I'm not sure I can trust them again. I'm investigating open weight models.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#403
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I see it more as a lose/lose: Any malicious user/attacker will just bypass the guardrails using one of a million established techniques for doing so while legit developers and security researchers will be prevented from finding problems by them.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#404
post #46

This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you , but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to…

Someone on here once point out that their CTO worked at Oracle and I haven't been able to forget that since.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#405

Earlier quoted context omitted.

This logic works only if distilling Claude is the only way to create another SOTA LLM, which is not the case.

How do you think the Qwen and MiniMax models perform so similarly to Anthropic frontier models? What is your take then?

Well Anthropic did not ask for permission before they distilled copyrighted material.

At least the Chinese have the decency of giving back the model weights and not put BS censorship because “it’s too dangerous”.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#408

I don't think they can convince me they have actually reversed course on this. Its invisible so we wouldn't know if they kept on doing it secretly. It required building out technical capability which is unlikely to remain forever unused while conveniently available to them. They relied on trust that they were providing the service they were being paid for. That trust was blown, and an "oops, lets undo that" does not…

Yes they already had an accident where the model magically downgrades itself, very likely that it just produces less good output rather than just stops working isn’t it… my guess is they were testing these features, accidentally or not, and wrote up something to justify what people were seeing. I find it absolutely disgraceful I can’t trust it to learn ML any more without there being a chance it’s messing me around. This whole saga represents a huge loss of trust for me in Anthropic.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#409
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#410
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

There is no middle ground to shadow bans while getting your hard earned cash. It is fraud/Nigerian scam
Post reply on HN