Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

51–60 of 206 posts

Re: Anthropic's Safety Superpower

#52
post #23

"they by extension think that only they should have final say over AI generally. When you further combine this realization with the company’s pronouncements about AI’s ability to conduct all economic activity, you realize that Anthropic’s leadership effectively wants to have power over everything and everyone." That might be one of the most important points in the post. Very troubling.

The problem is... what's the alternative?

It's questionable whether the current government can even unite the talent required for this project. Seizing it might just push all the talent to Europe or China.

The idea of open-sourcing something that falls into the "national security" category is clearly a non-starter unless there's more powerful, classified models that can outmatch them.

I think Anthropic has clearly demonstrated the most responsibility here: they've been crying for regulations, they were careful about Project Glasswing, and they've got comically over-sensitive filters around numerous topics.

Re: Anthropic's Safety Superpower

#53
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

> The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one.

But is that last part actually true though? Sure, there might be 600B+ models available for download and local inference if you have the hardware, but does the users who use Anthropic switch over to those even if they're available even as hosted models? Seems like some do, most don't, Anthropic and Claude remains very popular among the people who use LLMs, there is no denying that.

Re: Anthropic's Safety Superpower

#54
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

Re: Anthropic's Safety Superpower

#55
post #3

(reposted) As I understand it, ITAR regulations for export controls have just been applied to any form of Mythos. These are overseen by U.S. Departments of State and Commerce, and forbid foreign nationals from access to any form of Mythos, either within or outside the U.S. Only U.S. citizens and immigrants that are holders of a "green card" may now access Mythos. It appears that Anthropic does not have internal contr…

> As long as Anthropic is a U.S. company, there is no escaping this.

Reminds me of the RISC-V Foundation → RISC-V International move to Switzerland. Around the time some dumbass Republicans tried to impose export restrictions on a set of open, world-wide used specifications.

Pandora's box has been opened, and there's no closing it. Capable AI models will be everywhere.

Re: Anthropic's Safety Superpower

#56

> The entire Anthropic origin story is rooted in the founders’ belief that OpenAI wasn’t taking safety seriously enough; the company believes that only they can control AI, and that because they uniquely care about safety, they are justified in trying to control everyone else, up to and including the U.S. government. Anthropic believes they have the responsibility to guard their tools from mis-use. That is all. They…

it's a pretty ridiculous stretch to attribute them thinking that OpenAI wasn't taking safety seriously enough (which is, among other things, a little bit evident from the fact that they no longer have a safety team at all) into asserting that they want to control the US government.

Re: Anthropic's Safety Superpower

#57
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

I'm uneducated on how distillation works at more than a basic level so forgive me if this is a stupid question.

Isn't "distillation" of another provider's model exactly how these models got training date in the first place: Massive amounts of the written word + Prompt -> Answer. Why wouldn't distillation produce similar "reasoning" in the new model? It's just inputs and outputs.

Re: Anthropic's Safety Superpower

#58
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

I disagree. It is not the model alone. It needs a system which capitalizes on it. And this is very complex. Hardware, software, architecture - it takes a lot to get it right.

Try running the latest OS models on a normal Mac or PC. Claude Fable and Mythos are systems not just pure models.

And of course marketing. Don't believe the hype.

I think Claude is often times underwhelming. Security concerns are also a concern companies have a blond spot for. The really toughest pro security (Yes, pro! Totally different framing!) company I know is Google after all.

What I can companies advise to do is, really having more than just bug bounties but a professional hacker team that does nothing else but attacking them the whole day and night 24/7. This needs to be coordinated with the government otherwise you might sound an alarm and will be SWATed for doing good. And I would pay them huge sums since the risk and fallout warrant such a treatment, not the standard wage.

Hackers are the real deal, not AI. Proof: Hackers using AI.

Re: Anthropic's Safety Superpower

#59

> To that end, I can certainly buy the case that Fable/Mythos is in fact more capable when it comes to identifying and exploiting security issues This has been covered before: https://aisle.com/blog/ai-cybersecurity-after-mythos-the-jag... ( https://news.ycombinator.com/item?id=47732020 ) > Anthropic’s cautious roll-out was justified. The problem with publicly releasing models, however, is that guardrails can be jail…

[deleted]

Re: Anthropic's Safety Superpower

#60
post #43
post #33

Earlier quoted context omitted.

> Regulatory capture is the OpenAI and Anthropic end goal, for certain. it has to be, because the other way around - the government taking over parts or the whole thing - is inevitable if the trend holds.

the inevitable trend is that numbers will be free and nobody will control the whole thing ai-celebrities are just clinging to relevance like all the other celebrities out there

HN is the builder side of the conversation, and in my experience, few safety people congregate here.

The safety side of tech is a PTSD inducing shit show. Governments are more than happy to champion age verification laws, because parents, around the world, are clamoring for anything to pump the breaks on the social media experiment.

Society outside of HN is quite tired of Tech, and I despair of figuring out a way to make this clear to the commentariat.

Post reply on HN