Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

161–170 of 206 posts

Re: Anthropic's Safety Superpower

#161

> Here’s the thing about these safety justifications: I think they work because, to Anthropic, they aren’t justifications. The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside,…

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules. So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them…

I think the second the company starts to classify "competition" as mis-use... the whole "they're not committing harm" line sort of goes out the window.

Modern society is built on the idea that competition is required from companies, and we seem to be exiting that age into a new world of monolithic, monopolistic, mega-corps. Personally, I find that a real route to dystopia.

Where do you draw the line here?

What happens when your car stops working because you're driving a tesla, but you're working on EVs for Honda or Ford?

What happens when your macbook stops working, because you decided to commit to changes to ARM software, or RISC-V?

---

And before you dismiss those, this is literally what Anthropic is doing TODAY. Using their tools to develop competing tools is something they classify as mis-use, and shut you down for doing.

Personally, I just can't accept that as a valid moral stance. Wonderfully successful, abusive, and dystopian? Absolutely. Moral? FUCK NO.

When tools turn themselves off because the manufacturer has decided it doesn't like how you're using them... you're a slave with no autonomy.

Re: Anthropic's Safety Superpower

#162

> The entire Anthropic origin story is rooted in the founders’ belief that OpenAI wasn’t taking safety seriously enough; the company believes that only they can control AI, and that because they uniquely care about safety, they are justified in trying to control everyone else, up to and including the U.S. government. Anthropic believes they have the responsibility to guard their tools from mis-use. That is all. They…

I'm going to challenge this thought. I think assuming you have the ability to guard a tool (that you're "selling" for profit) from mis-use is the definition of "controlling behavior". It's the kind of ethically myopic take that can only really exist in this new digital age - where tools aren't actually sold, they're just digitally rented. The most telling part of the "control" narrative is that they happily classify…

I don't always consider ethics at all in logic, so I guess you can call it ethically myopic.

Installing safeguards to prevent a tool from being used for certain things is a perfectly natural and common thing to do when you are providing the tool as a service. For example, blocking VPNs and open proxies from accessing a free service if those are a major source of spam and abuse. Note that Anthropic never provided the model for offline use in a form that includes DRM -- they are simply safeguarding the service that provides access to the hosted model. The only ethical concern I see here is that some of their safeguards are ones I wouldn't personally agree with, and in a world where dependence on a model is expected it can become an issue if the model refuses to perform in some cases, etc., but that doesn't automatically mean the refusal itself is unethical unless that issue was known and expected (and unless the alternative is not bigger, worse bads)

Also note you are not even "digitally renting" anything. This is the exact same type of thing as, say, real humans in real life refusing to perform services for certain clients or under certain circumstances. Networking makes it possible to decouple some of these things, but that doesn't magically make it renting or automatically turn a refusal to perform services into an attempt to control clients. Just the same as I can choose to refuse any request, which does not automatically constitute attempted control over the asker. There can be ethical concerns about whether my refusal causes problems that I'm obligated to avoid (and whether or not such obligation exists), but that doesn't automatically contaminate the refusal itself unless I have knowledge of and intend the bad.

To use a much more relevant example, Anthropic's refusal to allow its models to be used for war (among other things) does not constitute any attempt to prevent war. It's only a refusal to assist in it. That's not some unfair, unethical attempt at controlling the government, that's just Anthropic not wanting to be responsible for assisting in war.

Re: Anthropic's Safety Superpower

#163

Earlier quoted context omitted.

"Capability per parameter" is rising, but parameter count remains an advantage. And small models remain bad, because "good" is a rapidly moving target. A 2026 4B beats 2024 4B, but both are far behind the contemporary frontier. Which makes them bad. There is no such thing as "too much capability" - a "good" model is whatever the current frontier is. In 2024, a "good" model is one that can be trusted to write a 800 li…

> A 2026 4B beats 2024 4B, but both are far behind the contemporary frontier. The thing about engineering is you don't just use the biggest bolt on the market on every bridge. > In 2024, a "good" model is one that can be trusted to write a 800 line script. In 2026, it's a model that can be trusted to do gnarly high-level planning and execution both This sounds a lot like having a single diamond-head hammer as the onl…

Good enough? That's a lie people tell each other because they lack imagination.

"It's good enough" was said about GPT-4, o1, o3, Opus 4 and more. Guess what happened? Newer models released, people updated their expectations of what LLMs can do, usage got more aggressive, and somehow, GPT-4 went from "good enough" to "obsolete trash".

If you have no imagination, then at least substitute your pattern recognition for it.

The world is hungry for capabilities. There are piles upon piles of tasks that aren't done by LLMs simply because LLMs aren't good enough to do them.

The thing a frontier model gives you is "you don't have to babysit a model to get it to do X", and that X gets more and more impressive release to release.

Re: Anthropic's Safety Superpower

#164

Earlier quoted context omitted.

I'm going to challenge this thought. I think assuming you have the ability to guard a tool (that you're "selling" for profit) from mis-use is the definition of "controlling behavior". It's the kind of ethically myopic take that can only really exist in this new digital age - where tools aren't actually sold, they're just digitally rented. The most telling part of the "control" narrative is that they happily classify…

I don't always consider ethics at all in logic, so I guess you can call it ethically myopic. Installing safeguards to prevent a tool from being used for certain things is a perfectly natural and common thing to do when you are providing the tool as a service. For example, blocking VPNs and open proxies from accessing a free service if those are a major source of spam and abuse. Note that Anthropic never provided the…

Where do you draw the line here?

What happens when your car stops working because you're driving a tesla, but you're working on EVs for Honda or Ford?

What happens when your macbook stops working, because you decided to commit to changes to ARM software, or RISC-V?

Tools should be neutral. The idea that a tool can only be wielded in a manner that its manufacturer approves of is... scary.

That's a real quick hop and a jump to a really, really ugly spot, societally speaking.

And sure - technically Anthropic is selling a service, but even that idea makes me quietly upset. The only reason they don't sell a product as a tool itself is that they have more control over the model as a service, and expect to be able to extract even more profit from their customers with this route.

---

My real hope is that open models FUCKING CRUSH them. Because almost nothing is scarier than a self-righteous, moral zealot.

Re: Anthropic's Safety Superpower

#165

Earlier quoted context omitted.

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules. So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them…

I think the second the company starts to classify "competition" as mis-use... the whole "they're not committing harm" line sort of goes out the window. Modern society is built on the idea that competition is required from companies, and we seem to be exiting that age into a new world of monolithic, monopolistic, mega-corps. Personally, I find that a real route to dystopia. Where do you draw the line here? What happen…

... you're perfectly welcome to compete with them, they're just not willing to provide their own product for that purpose. I also can't buy the schematics for Intel's latest chips, the source code for Windows, or a rocket from SpaceX to disassemble and study.

Like, there's a very critical point where you are asking to use their servers to directly compete with them.

This has been normal for somewhere between "decades" and "the entire history of commerce"

Re: Anthropic's Safety Superpower

#166

Earlier quoted context omitted.

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules. So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them…

Incomparable domains. People routinely suffer illness. We can compare outcomes. These ideologues are building something completely unprecedented which, according to themselves apparently, can go paperclip-rogue if one is not careful. So the worst case is unprecedented. Then there is the more mundane matter of heating up the economy, something which also has no one blameworthy until any such supposed bubble actually p…

No, I'm saying they have an actual track record that you can evaluate, and if you do so you will see that they have killed approximately zero people, etc. etc.

Or even just propose an alternative: who has a better track record, here? Who DO you trust?

Re: Anthropic's Safety Superpower

#167

Earlier quoted context omitted.

I don't always consider ethics at all in logic, so I guess you can call it ethically myopic. Installing safeguards to prevent a tool from being used for certain things is a perfectly natural and common thing to do when you are providing the tool as a service. For example, blocking VPNs and open proxies from accessing a free service if those are a major source of spam and abuse. Note that Anthropic never provided the…

Where do you draw the line here? What happens when your car stops working because you're driving a tesla, but you're working on EVs for Honda or Ford? What happens when your macbook stops working, because you decided to commit to changes to ARM software, or RISC-V? Tools should be neutral. The idea that a tool can only be wielded in a manner that its manufacturer approves of is... scary. That's a real quick hop and a…

> What happens when your car stops working because you're driving a tesla, but you're working on EVs for Honda or Ford?

> What happens when your macbook stops working, because you decided to commit to changes to ARM software, or RISC-V?

These sound similar, but aren't the same thing I'm talking about. It's more similar to a rideshare company refusing to serve you, or a cloud PC cutting your access.

It's not the same thing as DRM, which is when invasive malware attempts to control what you do with your devices -- your property -- and your resources, that you own. A car that can be shut off remotely, or that can detect competitive conditions and cease operation, is not the same as a hosted service refusing to have you. It is DRM.

Likewise, a MacBook that stops working based on your affiliations or your activity is not the same either. It is DRM. (Technically, Apple Activation Lock is DRM. So are locked Android bootloaders that can't be flashed with custom verified-boot signing keys, etc.)

When you're using someone else's private resources, someone else's private infrastructure, by default they have the right to simply no longer serve you, at any time.

DRM, by contrast, is when a machine or software you already own decides it will no longer function for you.

Yes, this is scary. We are already confronting this right now. Faceless corporations abruptly cut your access, or ban you for life from really important things. PayPal steals your money and pockets it instead of giving it back. Thousands, tens of thousands, hundreds of thousands of dollars (or whatever) gone because they said so. It's awful. It ruins lives.

But this happens because you need to draw a different line. Not whether one should be allowed to refuse services ever, but when. It's not okay that PayPal can steal tens or hundreds of thousands of dollars or more from you with no recourse, just because something looked suspicious, and you'll never know what it was even if you bankrupt yourself with arbitration costs (since their terms of service say you are simply not allowed to sue them anymore, and for some reason companies are allowed to just say this now and have it be legally binding).

A refusal at that point can be devastating, especially when it's for no fucking reason. Or when it's for a completely shitty and unjust reason, like in banning all accounts that have been involved in buying or selling adult content. There are cases where you can ruin someone's life by suddenly refusing to serve them with no recourse, and those are the cases that shouldn't be allowed to continue.

So push for that. Businesses shouldn't all have to serve every customer, but they shouldn't be able to just suddenly ruin someone's life. That's the scary part.

> The only reason they don't sell a product as a tool itself

...it's not the only reason. You can't exactly run trillion-parameter-scale models on the kinds of hardware people tend to already have in practice. And who's gonna pay tens of thousands of dollars up-front for their own inference hardware for the thing? (If you sold it as a hardware product.)

> My real hope is that open models FUCKING CRUSH them.

I would like this too!

Re: Anthropic's Safety Superpower

#168

Earlier quoted context omitted.

> A 2026 4B beats 2024 4B, but both are far behind the contemporary frontier. The thing about engineering is you don't just use the biggest bolt on the market on every bridge. > In 2024, a "good" model is one that can be trusted to write a 800 line script. In 2026, it's a model that can be trusted to do gnarly high-level planning and execution both This sounds a lot like having a single diamond-head hammer as the onl…

Good enough? That's a lie people tell each other because they lack imagination. "It's good enough" was said about GPT-4, o1, o3, Opus 4 and more. Guess what happened? Newer models released, people updated their expectations of what LLMs can do, usage got more aggressive, and somehow, GPT-4 went from "good enough" to "obsolete trash". If you have no imagination, then at least substitute your pattern recognition for it…

I wish you had addressed at least one of arguments in good faith before jumping to insults and countering a strawman argument I didn't make - I never claimed their will be no use for more capable models.

You do your AI-maximalism, and I'll stick to making trade-offs based on the needs of each piece of work.

Re: Anthropic's Safety Superpower

#169

Earlier quoted context omitted.

> does the users who use Anthropic switch over to those even if they're available even as hosted models? I'm currently spending $200 for Claude. That's around my maximum that I can afford. I could stretch that to $500 I guess. But I saw reports of people spending tens of thousands of dollars with Claude API. That's certainly outside of my budget. So if/when Anthropic decides to stop subsidizing subscription (if they…

The ai labs would be very dumb to get rid of subscriptions. First, I don’t even think the subscriptions are losing money, I suspect they’re around break even, maybe small loses. More importantly, the subscriptions are how they lock in users and convince companies to pay api rates. Without user loyalty that they cultivate with subscriptions businesses will just use the cheapest model on open router or maybe local mode…

> I don’t even think the subscriptions are losing money, I suspect they’re around break even, maybe small loses

whats the basis for this thought

Re: Anthropic's Safety Superpower

#170
"On one hand, I actually don’t begrudge Anthropic not wanting to help its competitors; on the other hand, what should be blisteringly clear is that Anthropic does not think that anyone else other than them should even be making frontier LLMs."

I don't find this blisteringly clear at all. A company making it harder for competitors to steal their IP is perfectly normal. This is Ben Thompson's personal grudge against Anthropic showing, yet again. He can't think rationally about this company.

Post reply on HN