Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

61–70 of 206 posts

Re: Anthropic's Safety Superpower

#61

> Here’s the thing about these safety justifications: I think they work because, to Anthropic, they aren’t justifications. The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside,…

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules.

So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them pulling a "greater good" card because they're in it for entirely selfish reasons.

Re: Anthropic's Safety Superpower

#62
post #27

Earlier quoted context omitted.

Could Anthropic relocate to a different country?

Individuals can leave, but the company cannot transfer restricted intellectual property. Europe has extradition treaties, so the U.S. can force anyone in Europe back to the U.S. for criminal indictment who demonstrates inappropriate possession of this technology.

Would be very hard to demonstrate that they did that. If all employees move to some country with a slow legal justice system and strong labor laws, they also recreate the training data because that can be transferred, they can train another version in said country which is perfectly legal.

Can you demonstrate beyond any reasonable doubt that the model weights have been transferred? No. Will the EU judges move to extradite said individuals (and many are EU citizens)? Also no, especially in the face of spurious accusations. And even if they were open to, you can stonewall everything and you will probably outlast any US administration pursuing that.

Re: Anthropic's Safety Superpower

#63
post #57

Earlier quoted context omitted.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

I'm uneducated on how distillation works at more than a basic level so forgive me if this is a stupid question. Isn't "distillation" of another provider's model exactly how these models got training date in the first place: Massive amounts of the written word + Prompt -> Answer. Why wouldn't distillation produce similar "reasoning" in the new model? It's just inputs and outputs.

Among other things, because you simply can't get those "massive amounts" of text from a SOTA model at reasonable cost. And complex reasoning cannot possibly be trained in a pure one-shot fashion, real post-training takes massive resources. The whole story doesn't add up.

Re: Anthropic's Safety Superpower

#64

Perhaps they should consider leaving the US. Pretty clearly the descent into a corrupt autocracy is having real consequences.

Oh please, the earlier spat with the Trump admin was the best thing that ever happened to Anthropic. Before that, Claude was really only well-known in developer circles, not the wider normie-sphere. After Anthropic got the "Trump hates them, so it MUST be good!" stamp of approval, the company's recognition and popularity took off.

This too, will end up being a good thing for them. The ban will end up getting lifted due to some "amazing deal" in the coming weeks and Anthropic will now have the "Trump tried to ban them, so they MUST have the most advanced AI model in the world!" stamp of approval just before IPO.

All this stuff is pro wrestling kayfabe.

Re: Anthropic's Safety Superpower

#65
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

This is totally inaccurate, the APIs provide the reasoning logs. You ABSOLUTELY can distill from APIs, in fact, that's the primary way distillation is done currently.

Re: Anthropic's Safety Superpower

#66
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

I disagree. It is not the model alone. It needs a system which capitalizes on it. And this is very complex. Hardware, software, architecture - it takes a lot to get it right. Try running the latest OS models on a normal Mac or PC. Claude Fable and Mythos are systems not just pure models. And of course marketing. Don't believe the hype. I think Claude is often times underwhelming. Security concerns are also a concern…

> I disagree. It is not the model alone. It needs a system which capitalizes on it. And this is very complex.

AFAICT … despite saying you “disagree”, you appear to be agreeing with the parent comment that the model is less important and compute (all that complex infra) and data (also complex infra) are more important.

Re: Anthropic's Safety Superpower

#67
post #43

Earlier quoted context omitted.

the inevitable trend is that numbers will be free and nobody will control the whole thing ai-celebrities are just clinging to relevance like all the other celebrities out there

HN is the builder side of the conversation, and in my experience, few safety people congregate here. The safety side of tech is a PTSD inducing shit show. Governments are more than happy to champion age verification laws, because parents, around the world, are clamoring for anything to pump the breaks on the social media experiment. Society outside of HN is quite tired of Tech, and I despair of figuring out a way to…

> Society outside of HN is quite tired of Tech, and I despair of figuring out a way to make this clear to the commentariat.

I don't think anyone in tech is really truly engaging with how quickly the shine has come off the tech industry. Except maybe Apple, who even so still have some work to do.

Re: Anthropic's Safety Superpower

#68

Earlier quoted context omitted.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

This is totally inaccurate, the APIs provide the reasoning logs. You ABSOLUTELY can distill from APIs, in fact, that's the primary way distillation is done currently.

Not for proprietary models, all you get is a terse summary.

Re: Anthropic's Safety Superpower

#69
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

Do you think token completion endpoints are the final form for AI APIs?

Re: Anthropic's Safety Superpower

#70
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

I disagree. It is not the model alone. It needs a system which capitalizes on it. And this is very complex. Hardware, software, architecture - it takes a lot to get it right. Try running the latest OS models on a normal Mac or PC. Claude Fable and Mythos are systems not just pure models. And of course marketing. Don't believe the hype. I think Claude is often times underwhelming. Security concerns are also a concern…

For now I suspect however that the gigantic models are not needed and you will be able to do pretty much what you need in a specific domain with 120b or lower. There is so much trash in the frontier models. I don't need all the world's slam poetry for my coding tasks for example.
Post reply on HN