Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

81–90 of 206 posts

Re: Anthropic's Safety Superpower

#81
Safety is a cost center, the internal team who sends you the bills when you move fast and break things.

I always thought safety was interesting in and of itself, but for some reason HN doesn’t have many people from the safety side of tech in conversation.

Tech isn’t a niche hobby anymore; Billions of people are impacted by the decisions of a few firms.

My grandfathers android had 3 different messaging apps installed, somehow. AI is enabling new forms of fraud at a time when we still haven't solved the old ones.

And this is all in the first world, move your coordinates to the developing world? We had human trafficking to get educated English speakers into call centers in Laos/Cambodia to defraud first world inhabitants of their money.

We aren’t in the early days of tech anymore, and the kind of scale that we have enabled comes with it a certain cost. We can choose to ignore them, or to understand them, but we will feel their impacts all the same.

Re: Anthropic's Safety Superpower

#82
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

> The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. But is that last part actually true though? Sure, there might be 600B+ models available for download and local inference if you have the hardware, but does the users who use Anthropic switch over to those even if they're available even as hosted models? Seems l…

> does the users who use Anthropic switch over to those even if they're available even as hosted models?

I'm currently spending $200 for Claude. That's around my maximum that I can afford. I could stretch that to $500 I guess. But I saw reports of people spending tens of thousands of dollars with Claude API. That's certainly outside of my budget.

So if/when Anthropic decides to stop subsidizing subscription (if they ever do that thing, I still not sure about that), I'll certainly look at the other options. And available "open weights" LLMs hosted by someone will be my first pick. Right now Claude 4.8 feels very advanced, but things move very fast...

Re: Anthropic's Safety Superpower

#83
post #67

Earlier quoted context omitted.

HN is the builder side of the conversation, and in my experience, few safety people congregate here. The safety side of tech is a PTSD inducing shit show. Governments are more than happy to champion age verification laws, because parents, around the world, are clamoring for anything to pump the breaks on the social media experiment. Society outside of HN is quite tired of Tech, and I despair of figuring out a way to…

> Society outside of HN is quite tired of Tech, and I despair of figuring out a way to make this clear to the commentariat. I don't think anyone in tech is really truly engaging with how quickly the shine has come off the tech industry. Except maybe Apple, who even so still have some work to do.

Technology and science is the intersection that is supposed to make our lives better, easier, more prosperous. The last decade or two what marvelous technology has came from silicon valley that hasn't served primarily the billionaire class and made life worse for the common people.

The yoke of silicon valley is feeling heavy. People might just throw it off.

Re: Anthropic's Safety Superpower

#84
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

I struggle with the practicality of the whole thing.

The amount of tokens required to properly distill a frontier model is so large that by the time you could consume the # of tokens you would either be banned for extremely obvious abuse or a new model would be released, rendering your efforts less and less valuable over time. Intelligence is not a linear thing. Being behind just a little bit can have exponential consequences.

Re: Anthropic's Safety Superpower

#85
post #23

"they by extension think that only they should have final say over AI generally. When you further combine this realization with the company’s pronouncements about AI’s ability to conduct all economic activity, you realize that Anthropic’s leadership effectively wants to have power over everything and everyone." That might be one of the most important points in the post. Very troubling.

The problem is... what's the alternative? It's questionable whether the current government can even unite the talent required for this project. Seizing it might just push all the talent to Europe or China. The idea of open-sourcing something that falls into the "national security" category is clearly a non-starter unless there's more powerful, classified models that can outmatch them. I think Anthropic has clearly de…

I think the overly sensitive filters reveals an alignment gap

if they had more success on alignment and safety research then I don't think the cludgy filters would be necessary

Re: Anthropic's Safety Superpower

#86
post #77
post #32

Earlier quoted context omitted.

what Dario wants is to retain any influence whatsover on how the research progresses before the inevitable nationalization of the frontier. he gets to keep the N-2 tech and maybe influence the N-1 tech, but the only influence on the frontier he has is today; whatever he imprints in the pipeline the government takes over. IOW I don't think he thinks in the same categories as most folks here.

N-1? N-2?

Best-possible-model (N) - Two Generations (2), same with N-1, N is the SOTA in this example. I'm not sure that actually clarifies what the comment is trying to say other than they think the models will be nationalized (can't even imagine what that would look like).

Re: Anthropic's Safety Superpower

#87
post #50
post #42

Earlier quoted context omitted.

the whole thing playing out as expected. if you think about it, the only question is the timeline. the next model with a gap to mythos as mythos is to opus will be controlled technology from the get-go. the one after it may be top secret.

Open models will catch up eventually, TOTL models will get distilled into smaller, more efficient versions, it’s not something you can moat indefinitely

who is going to continue to publish these open models and why would they keep doing it?

Re: Anthropic's Safety Superpower

#88
post #57

Earlier quoted context omitted.

"Distillation" from APIs is not a thing, it cannot replicate a model's deep reasoning and behavior.

I'm uneducated on how distillation works at more than a basic level so forgive me if this is a stupid question. Isn't "distillation" of another provider's model exactly how these models got training date in the first place: Massive amounts of the written word + Prompt -> Answer. Why wouldn't distillation produce similar "reasoning" in the new model? It's just inputs and outputs.

What you're describing is (pre-)training. Distillation requires richer labels, the probability distribution over tokens (it would be logits rather than probabilities but that's not important). From a chat transcript you can only understand the argmax/most likely token of that distribution (and only if the API allows you to set the temperature to 0). It's not impossible for an API to give you that but they won't if they don't want you distilling their models.

The intuition is that distillation exploits not only the "right" answer but the relationship between answers (what's the second most right answer? the third? etc).

Re: Anthropic's Safety Superpower

#89
post #42

Earlier quoted context omitted.

the whole thing playing out as expected. if you think about it, the only question is the timeline. the next model with a gap to mythos as mythos is to opus will be controlled technology from the get-go. the one after it may be top secret.

Or OpenAI will pay Trump's regime's bribe and they'll suddenly realise that it does not need controlling and they're free to sell it?

that's... an optimistic take I think

Re: Anthropic's Safety Superpower

#90

> Here’s the thing about these safety justifications: I think they work because, to Anthropic, they aren’t justifications. The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside,…

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules. So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them…

Incomparable domains. People routinely suffer illness. We can compare outcomes. These ideologues are building something completely unprecedented which, according to themselves apparently, can go paperclip-rogue if one is not careful. So the worst case is unprecedented. Then there is the more mundane matter of heating up the economy, something which also has no one blameworthy until any such supposed bubble actually pops.

> So... what, you just don't trust anyone good?

The baseline here is apparently that they are good, I’m just supposed to trust and shut up?

Post reply on HN