Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

311–320 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#311
post #15

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I think it's a big mistake to conflate the cyber (and bio) refusals with the LLM development refusals.

I can sympathize with the argument for the cyber refusals - especially as a temporary measure - especially if Mythos is available to those trying to defend against vulnerabilities.

The LLM development nerfing (and now refusals) is very different though. Anthropic has even said it isn't just for safety reasons:

> Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms.

It's at least partially an anti-competitive measure.

The closest analogy is putting measures in a compiler to stop it being able to build other compilers.

Another analogy is priesthoods with secret religious knowledge that "only they are qualified to know".

Re: Anthropic apologizes for invisible Claude Fable guardrails

#312

Earlier quoted context omitted.

This logic works only if distilling Claude is the only way to create another SOTA LLM, which is not the case.

How do you think the Qwen and MiniMax models perform so similarly to Anthropic frontier models? What is your take then?

Probably the same reason a Epyc 9965 from hetzner performs just as well as one from AWS for one tenth the cost.

Anthropic is offering a commodity product and trying to convince you it isn’t.

It’s even in the name, it’s a myth and a fable. Never happened doesn’t exist.

Also I believe at least on coding that qwen is now the frontier model, fable is its copy of frontier models. In the same way that the Ferrari Luce is an expensive imitation of a SU7 Ultra.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#315

Earlier quoted context omitted.

No. You read the actual essay, then explain how we're supposed to interpret this more charitably: Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety. I am grateful to see the Trump administration’s Executive Order move incrementally towards a great…

You got baited by a confirmed Anthropic shill, see more info here: https://news.ycombinator.com/item?id=48270186

Confirmed by you!

I don't really agree with their point here, but there are plenty of people in the AI community whose views are aligned with Anthropic's. That doesn't make them shills.

It's actually important those views are put forward.

A place like LessWrong has the opposite problem - there is no one there who questions the "safety narrative" so the discussion swings more and more towards the extreme end of that spectrum.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#317
Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right?

Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Re: Anthropic apologizes for invisible Claude Fable guardrails

#318
post #153

Earlier quoted context omitted.

Even with them making those guardrails visible, it's a bit ridiculous in my eyes. I have been experimenting with smaller models, will Claude assume I'm some Chinese or Russian agent trying to distill their secrets and bar me from learning? Because that's insane. What if I discover a more efficient way to build models with Claude? Well, we'll never know now. What if someone else entirely could discover a breakthrough…

The whole shtick is to get you addicted whilst reducing your ability to go without, acquire power over you, jack up the prices whilst manipulating the quality of the tokens/output available to you. Cant believe how stupid people are. You couldnt see this coming? Shame on you.

I already made up my mind, I'm not using that model if its sending proprietary code over to Anthropic, they can kiss my rear. If every frontier model winds up doing this, I will stop using them. There's plenty of employers / jobs where this is not okay behavior from an LLM.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#320

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Not really, the purpose of Excel is pretty clear cut and the scope is small.

Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.

Post reply on HN