Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

321–330 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#321

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

That analogy is... Not inappropriate, but I think it could confuse by being compatible with two different problems, where only one is the target of today's controversy.

1. The sloppy/unpredictable behavior of LLMs as a general class of algorithm, how you shouldn't use document-generation for calculating budgets, and you shouldn't trust it to not-alter things you "asked" it to to alter.

2. Vendors of thing-as-a-service (not necessarily only LLMs) putting in traps and sabotage to prioritize their own business-model or economic incentives.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#322
post #320

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.

What’s the point when they will remove those guardrails when competition reaches their levels. Shows that they don’t Reddit care about “safety” at all

Re: Anthropic apologizes for invisible Claude Fable guardrails

#324
post #320

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.

No. Excel is a general purpose tool that can be used for calculating tasks that are good, neutral, or evil things. It's a fancy calculator.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#325
post #311
post #15

Earlier quoted context omitted.

I think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.

I think it's a big mistake to conflate the cyber (and bio) refusals with the LLM development refusals. I can sympathize with the argument for the cyber refusals - especially as a temporary measure - especially if Mythos is available to those trying to defend against vulnerabilities. The LLM development nerfing (and now refusals) is very different though. Anthropic has even said it isn't just for safety reasons: > Usi…

The Anthropic refusal description is even more direct.

“The request could assist the development of competing AI models, which is restricted under Anthropic's commercial terms. Benign machine learning work can also trigger this category.”

Source: https://platform.claude.com/docs/en/build-with-claude/refusa...

Re: Anthropic apologizes for invisible Claude Fable guardrails

#326

Earlier quoted context omitted.

Distillation is not a thing unless you actually have the model weights. What people misleadingly call distillation is just training on chat logs, which has always been routine practice in the industry. There's a reason why every model today talks like early releases of ChatGPT.

If Anthropic is calling it distillation [1] then that would argue for it being correct (or at least canonical) terminology. [1] https://www.anthropic.com/news/detecting-and-preventing-dist...

No, a company choosing to use some terminology doesn’t make it correct nor canonical in any sense; especially when they have a vested interest in not being neutral or credible.

If Google starts calling ads “Best Links” that doesn’t make it correct nor canonical; the correct term is still ads.

Traditionally, distillation is when you get the actual logits of a model response (not exposed via API for years) and then use that to train a model.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#327

I like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thin…

> paternalism isn't a good look.

Anthropic doesn't care. The goal right now is simply to avoid any and all bad PR on the way to the cashout IPO.

And paternalism will generate far less bad PR than somebody using AI on something that does real damage and makes headline news.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#328
post #269

Earlier quoted context omitted.

Yeah, I didn't mean to downplay how hard it is to apply the utilitarian calculus or even to suppose that the bare doctrine of utilitarianism resolves questions about what the ultimate good we should be trying to maximize is. I basically agree that utilitarianism is not a complete recipe for how to live. I just think that it probably gives the correct answer in cases where we can see clearly how to apply it because I'…

This kind of reasoning leads you to reasoning that if he was an ineffective fraudster, it would be less moral, as he would have bought less mosquito nets. So it’s not only moral to do fraud, but you most extremely competently do fraud. I think this being a reasonable utilitarian point to make is not a point in utilitarianism’s favor.

This point is very similar to the core plot of Watchmen

Re: Anthropic apologizes for invisible Claude Fable guardrails

#329

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

They implemented both those things, but only apologized for the first. They’re doubling down on the second.

My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes.

I’m definitely shopping around for other LLM providers next week, and testing vs local (target: 128GB strix halo - any war stories?)

Re: Anthropic apologizes for invisible Claude Fable guardrails

#330
post #312

Earlier quoted context omitted.

How do you think the Qwen and MiniMax models perform so similarly to Anthropic frontier models? What is your take then?

Probably the same reason a Epyc 9965 from hetzner performs just as well as one from AWS for one tenth the cost. Anthropic is offering a commodity product and trying to convince you it isn’t. It’s even in the name, it’s a myth and a fable. Never happened doesn’t exist. Also I believe at least on coding that qwen is now the frontier model, fable is its copy of frontier models. In the same way that the Ferrari Luce is a…

China no. 1?
Post reply on HN