Earlier quoted context omitted.
Then what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"
They are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.
Anthropic apologizes for invisible Claude Fable guardrails
411–420 of 489 posts
Re: Anthropic apologizes for invisible Claude Fable guardrails
#412Earlier quoted context omitted.
public safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.
Distillation is not a thing unless you actually have the model weights. What people misleadingly call distillation is just training on chat logs, which has always been routine practice in the industry. There's a reason why every model today talks like early releases of ChatGPT.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#413Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
They implemented both those things, but only apologized for the first. They’re doubling down on the second. My limited experience with fable over the last few days suggests (1) I can’t see any improvement in output, and (2) it is useless for writing secure software because it constantly hits safety walls if you ask it to close security holes. I’m definitely shopping around for other LLM providers next week, and testi…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#414Re: Anthropic apologizes for invisible Claude Fable guardrails
#415Earlier quoted context omitted.
What is "EA" in this context? I see a lot of people using this initialism.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
I understand how one may wonder if there was a way to do that, but it feels insane to me that one would actually conclude that "yes, it is possible". We have examples everywhere showing that it is generally impossible to define a metric that correctly represents the underlying concept we want to measure.
Said differently, I feel like Effective altruism fundamentally starts by saying "I don't believe in Goodhart's law". Which seems intellectually dishonest to me.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#416Can you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
Not really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.
That has to be the understatement of the century.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#417Re: Anthropic apologizes for invisible Claude Fable guardrails
#418Anthropic seems to keep making the same mistake. Not being upfront or direct about random things, that come back and bite them. It isn't exactly unethical. Perhaps, ethically incompetent.
Re: Anthropic apologizes for invisible Claude Fable guardrails
#419Earlier quoted context omitted.
Guess FTX disproved the concept of giving to effective charities, time to start donating to my church again.
todays EA is not about giving to charities, that was the original mission with 40k hours and ethereum (i think vitalik still believes in this version). then the yudkowsky xrisk/ai safety crowd took over lesswrong and turned it into a cult. now its utilitarianism taken to the extreme. if you believe a skynet scenario killing everyone on earth is plausible then the "logical" thing to do is allow literally anything in t…
Re: Anthropic apologizes for invisible Claude Fable guardrails
#420Earlier quoted context omitted.
Effective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism
If you ban women from driving you can eliminate around half the car accidents. Don't you want to reduce car related deaths??