Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

41–50 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#41

Earlier quoted context omitted.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

They performed famously well at FTX.

Guess FTX disproved the concept of giving to effective charities, time to start donating to my church again.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#43
post #36

The idea of them purposefully wasting my time by having the model act dumber and me having to argue with it without knowing if it’s the prompt or the model was just such an idiotic product decision I can’t believe they shipped that without getting any feedback from users first.

[flagged]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#44

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look.

In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt.

Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#45

Earlier quoted context omitted.

They didn't get "caught." It was published, by them, when they released Fable a few days ago. They were very clear about it. It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.

No, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798

Excusing? Their comment is factually correct and the parent is factually wrong.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#46
This has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you, but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to vibe code some dashboards, a web app or let it drive Excel, but anything more interesting than that is forbidden.

If it was just plain monetary concerns and sabotage of competitors I'd almost be fine with it, but it seems they actively want to monopolize most of human progress in their enlightened hands, lest the mob does something undesirable with these powers.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#47

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO?

Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

Re: Anthropic apologizes for invisible Claude Fable guardrails

#48

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

Basically all critiques of Anthropic's policy moves on these topics boil down to people not believing the fundamental concerns are real, and often then going a step further to conclude that Anthropic doesn't actually believe their concerns either.

If you believe Anthropic believes what they say they do, all of it makes sense.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#50

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.

You are arguing with a straw man. Most are saying they should be explicit with the failure modes rather than fail silently. They aren't saying there should be no guardrails.
Post reply on HN