Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

71–80 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#71
post #55

Earlier quoted context omitted.

Basically all critiques of Anthropic's policy moves on these topics boil down to people not believing the fundamental concerns are real, and often then going a step further to conclude that Anthropic doesn't actually believe their concerns either. If you believe Anthropic believes what they say they do, all of it makes sense.

What are you referring to? The cult belief that they are ushering in a machine god or that they strictly care about making as much money as humanely possibly while ignoring the absolutely destructive impacts these companies have had on society? IMO they are using the cult messaging to distract the public so they take out all the oxygen in the room regarding people that care about the immediate impacts (climate exacer…

"Why don't they just not participate in the arms race?!" - guy who's never heard of arms races

If they believe they're creating "a machine god" and that it's better it's their machine god than someone else's (which, given the other contenders, I tend to agree with), then all the corollaries you mention are mostly irrelevant.

Whether you believe they're creating a machine god is irrelevant. They believe that they are. It would be helpful if you could create an actually good argument for why they cannot or are not creating a machine god, but it turns out there are no good arguments for why it's impossible to do so. And so... they shall try.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#72
post #15

Earlier quoted context omitted.

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

Claude Opus 4.6 and 4.8 find vulns in source code just fine and 4.6 will pentest without source for you given a proper harness WITHOUT jailbreaking. WITH jailbreaks, you can probably imagine what they are capable of. Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.

public safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#73

Earlier quoted context omitted.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

> Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Let's just assume it was " only " that? It's unreasonable to assume they are aiming to upset people who are just giving them money in the way they want. It makes no business sense, for any company. So that has to be a byproduct. Model training is one of the more expensive undertakings in the world right now…

It's about how they took measures against it. Sabotaging the requests is super shady and breaks all other areas of trust in the company their models.

All they had to do was have a simple, transparent output "Sorry, that request is against our terms of service. This session has been terminated"

Re: Anthropic apologizes for invisible Claude Fable guardrails

#74

Earlier quoted context omitted.

They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.

> guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors All this longtermism though is harmful. There are real problems of data theft, bias, labor displacement, and environmental costs that are happening right now but every push for regulation and regulatory capture, and all the safety talk, is always focused on some speculative futur…

I cannot overstate how much I think this take is wrong. Please please reconsider, look at the rate of progress being made, and consider that even if you only think ASI 'may' never happen in your lifetime it should still be one of your #1 concerns.

Honestly, that respect for 'copyright protections' has somehow become a leftist shibboleth is bizarre to me and indicative that something has become deeply warped in our discussions around this topic.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#75
post #68

Earlier quoted context omitted.

[flagged]

I think you can sympathize with the safety motives while still thinking this was a dumb implementation to degrade silently? I actually have faith in them getting the guardrail triggers pretty good, but consensus seems like they’re not yet there yet.

I think it is clear given the stakes why you would not want to make your guardrails probe-able/invertable.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#76

Earlier quoted context omitted.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

> Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Let's just assume it was " only " that? It's unreasonable to assume they are aiming to upset people who are just giving them money in the way they want. It makes no business sense, for any company. So that has to be a byproduct. Model training is one of the more expensive undertakings in the world right now…

The hidden safeguard was not against distilling, it was against "frontier" ML research with no indication whatsoever of what "frontier" might mean, but possibly even including research into model safety or alignment. That amounts to deliberately boobytrapping research across an entire legit academic field, which is ridiculously unaligned behavior.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#77
post #36

The idea of them purposefully wasting my time by having the model act dumber and me having to argue with it without knowing if it’s the prompt or the model was just such an idiotic product decision I can’t believe they shipped that without getting any feedback from users first.

[flagged]

We do understand why they did it, and the reason is dark and cynical.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#78

Earlier quoted context omitted.

That would be Anthropic.

Well, Anthropic thinks it should be the Trump administration [1]. This whole business just keeps getting dumber. 1: https://darioamodei.com/post/policy-on-the-ai-exponential

Read the actual essay. I cannot possibly imagine how you come to that conclusion unless you're just arguing in bad faith.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#79

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

[deleted]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#80

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

We've all been observing it. The recent spate of cyberexploits were powered by AI.
Post reply on HN