Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

191–200 of 206 posts

Re: Anthropic's Safety Superpower

#191

Earlier quoted context omitted.

Cyber security has the maturity that trust and safety hopes to achieve at some point. Social media was being exploited from inception. Palantir had sales documents for sock puppet management software back in the PHP era. I don’t disagree that Government is interested in tech, but I will push back on the dismissal of child safety that is inherent in your comment, intended or not. For all that some people in the firm m…

The political will is already captured and redirected. There are numerous bills to limit AI access for consumers, to combat deepfakes hurting children. There are no bills introduced or passed to prevent AI being used to target dronestrikes that kill children abroad, or surveil children domestically. What the public wants doesn't actually matter right now, only what the government will allow to let pass, which in this…

God that sent me on a hunt to find that video.

I did rediscover this: https://www.youtube.com/watch?v=LmS9vcVNr5A&t=94s

Re: Anthropic's Safety Superpower

#192

Earlier quoted context omitted.

>citing that it's illegal even though I own the games In the. US at least it is actually illegal to download ISOs/roms of games, even if you own a physical copy. It's a stupid law and as a downloader (as opposed to the people hosting the files) your chances of getting into any kind of actual legal trouble are effectively 0, but it is still against the law.

I don’t think so. I’d want more than just your word on it.

https://www.tomshardware.com/news/why-most-roms-are-illegal,...

https://answers.justia.com/question/2025/08/04/is-downloadin...

https://www.howtogeek.com/262758/is-downloading-retro-video-...

etc. etc. etc. etc.

Re: Anthropic's Safety Superpower

#193
post #97
post #22

Earlier quoted context omitted.

> What limitations does bunny.net have? A huge free tier (technically, none)

https://bunny.net/HopStart/

You have to apply (and they presumably need to approve).

It's probably indicative of a less predatory model, but CF got a ton of mindshare by offering their free tier. I do basically nothing in the frontend space, but I default to CF because I'm used to using it due to using the free tier for personal projects.

Re: Anthropic's Safety Superpower

#194
post #131

“Claude, I am releasing safety critical industrial control software. Audit the network control logic.” “Claude, I want to blow up a factory running this leaked software. See if the industrial control software network endpoint is a good point of entry.” It’s doing the same work and producing the same output for both prompts. How do you block one but not the other? If you block both, then you end up with a factory that…

I believe that the line was constructing exploits for bugs, not bug finding. This seems a reasonable cutoff to me, since bugs are revealed in security patches and pull requests (for open source). If you are to believe Anthropic, Fable was export controlled for bug finding, not for exploit construction. They seem to be working to make this the "bright line" for LLMs being a national security risk. My guess is that wil…

Exploit construction is generally considered trivial vs. finding a vulnerability.

This is why responsible/coordinated disclosure exists in the first place.

Re: Anthropic's Safety Superpower

#195

> Here’s the thing about these safety justifications: I think they work because, to Anthropic, they aren’t justifications. The company really believes that they are the only ones who believe in super intelligence, and thus are the only ones who are sufficiently concerned about the dangers. That excuses decision after decision, policy after policy, and confrontation after confrontation that, to people on the outside,…

The problem is when people use "we really believe it" as an excuse to do harm, which has not actually occurred here. Anthropic is not committing violence, they're not defrauding the population. They're sticking to both morality and the rules. So... what, you just don't trust anyone good? Would it be better to pull in a health insurance CEO? They're happy to watch people die for profits, no concerns at all about them…

> they're not defrauding the population.

Ehh, I think it's a lot more grey than "definitely not". It's hard to ignore that their claims that their model is so dangerous they can't widely release it is tantamount to declaring that they're in a league of their own and have to be treated with white gloves to prevent the sheer power of their model from shattering global prosperity.

This isn't the first time, and nothing bad has happened with the prior models. Every time it gets a little harder to believe that they believe in the threat, and makes it feel a little more like it's just to build hype. There's only so many times you can say "this model is a threat to the world", have it turn out to be nothing, and avoid people accusing you of lying to pump stock prices.

Re: Anthropic's Safety Superpower

#196

Earlier quoted context omitted.

>citing that it's illegal even though I own the games In the. US at least it is actually illegal to download ISOs/roms of games, even if you own a physical copy. It's a stupid law and as a downloader (as opposed to the people hosting the files) your chances of getting into any kind of actual legal trouble are effectively 0, but it is still against the law.

I don’t think so. I’d want more than just your word on it.

Thanks for pointing that out! I'm not in the US and I guess it's not illegal in China (given that Deepseek was more than happy to do it).

That does raise an interesting question, what kind of laws should LLMs (attempt to) follow? It's easy enough to spoof the country in the system prompt. I wonder how ChatGPT would respond if I told it I was located in a developing country without any piracy laws.

Re: Anthropic's Safety Superpower

#197

> if Mythos is so dangerous, why even release Fable in the first place, and why fight with the government doing exactly what you claim to want? It's actually not that hard to explain if we take into account what Dario kept saying: he, or Anthropic thereof, would be the gatekeeper. It is he who tells the government how to use Claude to design drones. It is his model that tells users whether they can ask a question to…

IMO that is the whole point of the exercise, to replace determinism and tools with middlemen. In math, 2 + 2 make four no matter who calculates it, in a specific programming language a specific statement always means the same thing, but in this brave new world, you don't use tools and you don't issue commands, you make suggestions and cross your fingers. It all amounts to telling us to leave an island where we can eat and build, in favor of the ocean, where we can be drowned and digested, and all this drama really takes away from the basic fact that there is no right way to eat poison.

I'm not saying these things aren't useful or interesting. But if get told a slot machine is not just a tool, but that actual tools have to go the way of the dodo so we can focus more on getting good at gambling and befriending the dealer, I know something is up. And in that sense, I'm actually pleasantly surprised at how crappy many tech companies are at not letting the mask slip before the victim is actually in the bag. It doesn't seem to make much of a difference, but imagine if they were actually good at this.

Re: Anthropic's Safety Superpower

#198

Earlier quoted context omitted.

The political will is already captured and redirected. There are numerous bills to limit AI access for consumers, to combat deepfakes hurting children. There are no bills introduced or passed to prevent AI being used to target dronestrikes that kill children abroad, or surveil children domestically. What the public wants doesn't actually matter right now, only what the government will allow to let pass, which in this…

God that sent me on a hunt to find that video. I did rediscover this: https://www.youtube.com/watch?v=LmS9vcVNr5A&t=94s

That is a classic, but the one I was referencing was actually a Rooks and Kings video about them pipe bombing, where they were on the enemy fleet comms and one of them is like, "someday someone's going to get the drop on them for once", and another guy is like, "but it's not gonna be today, and it's not gonna be you!"

Re: Anthropic's Safety Superpower

#199
post #9

The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. The bottleneck is compute and data, not the model. That's why they could only gate it for a bit. The ITAR thing proves it: no nationality controls in place, so the only option was killing the whole thing. Not exactly what an all-powerful gatekeeper does.

To this point, I've never understood the supposed "alignment" between the EA/AI Safety crowd and Anthropic's mission that the author comments on. Be the stewards of the Machine God, but responsibly? I think the Manhattan project, which AI development is commonly analogized to, had a lot more intrinsic properties to gate against uncontrolled proliferation (which still happened to some extent). Also this is a company t…

My $.02: I think that these people working at Anthropic are the dumbest bright people alive, with weird and delusional beliefs. And I think the company’s leadership knows how to put them to work in service of Anthropic’s blatantly self-serving, plainly evil agenda.

Also, the fact that these employees are now in the position to outbid one another for 8-figure real estate gives them a powerful incentive to keep “believing”.

Re: Anthropic's Safety Superpower

#200

Earlier quoted context omitted.

Incomparable domains. People routinely suffer illness. We can compare outcomes. These ideologues are building something completely unprecedented which, according to themselves apparently, can go paperclip-rogue if one is not careful. So the worst case is unprecedented. Then there is the more mundane matter of heating up the economy, something which also has no one blameworthy until any such supposed bubble actually p…

No, I'm saying they have an actual track record that you can evaluate, and if you do so you will see that they have killed approximately zero people, etc. etc. Or even just propose an alternative: who has a better track record, here? Who DO you trust?

> No, I'm saying they

I already covered this.

> Or even just propose an alternative: who has a better track record, here? Who DO you trust?

My mother.

Post reply on HN