Live data from Hacker News

Anthropic apologizes for invisible Claude Fable guardrails

theverge.com

411–420 of 489 posts

Re: Anthropic apologizes for invisible Claude Fable guardrails

#411

Earlier quoted context omitted.

Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"

They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.

Its not difficult to understand, they where fine with Claude being used to plan the murder of people around the world for Trump's war on peace, and the spying on any person as long as they are not from the USA. I don't see this any different then a company that makes missiles at this point, the blood is on their hands already.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#412

Earlier quoted context omitted.

public safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.

Distillation is not a thing unless you actually have the model weights. What people misleadingly call distillation is just training on chat logs, which has always been routine practice in the industry. There's a reason why every model today talks like early releases of ChatGPT.

You can logit distill (full token probabilities) or one hot distill (chat logs), or even align hidden states. All are distillation methods.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#413
post #329

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…

[dead]

Re: Anthropic apologizes for invisible Claude Fable guardrails

#415
post #10

Earlier quoted context omitted.

What is "EA" in this context? I see a lot of people using this initialism.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

I may be naive, but I have the feeling that "I will arbitrarily set numbers on things and call it impartial" is... weird at best.

I understand how one may wonder if there was a way to do that, but it feels insane to me that one would actually conclude that "yes, it is possible". We have examples everywhere showing that it is generally impossible to define a metric that correctly represents the underlying concept we want to measure.

Said differently, I feel like Effective altruism fundamentally starts by saying "I don't believe in Goodhart's law". Which seems intellectually dishonest to me.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#416
post #320

Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?

Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.

> the purpose of Excel is pretty clear cut and the scope is small.

That has to be the understatement of the century.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#417
I find it interesting that when a government tries to "put guardrails" (whatever they try) they are immediately considered authoritarians, but when a private company that has waay too much power for an entity that is not elected does that, people seem much less opposed.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#418

Anthropic seems to keep making the same mistake. Not being upfront or direct about random things, that come back and bite them. It isn't exactly unethical. Perhaps, ethically incompetent.

It’s because they are themselves deluded by their marketing story about their own product.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#419
post #159

Earlier quoted context omitted.

Guess FTX disproved the concept of giving to effective charities, time to start donating to my church again.

todays EA is not about giving to charities, that was the original mission with 40k hours and ethereum (i think vitalik still believes in this version). then the yudkowsky xrisk/ai safety crowd took over lesswrong and turned it into a cult. now its utilitarianism taken to the extreme. if you believe a skynet scenario killing everyone on earth is plausible then the "logical" thing to do is allow literally anything in t…

As 8note is pointing out, Eliezer Yudkowsky didn't "take over" Less Wrong, he founded it.

Re: Anthropic apologizes for invisible Claude Fable guardrails

#420
post #187

Earlier quoted context omitted.

Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism

If you ban women from driving you can eliminate around half the car accidents. Don't you want to reduce car related deaths??

Banning white people would reduce it by a much greater amount, at least in North America.
Post reply on HN