Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

121–130 of 206 posts

Re: Anthropic's Safety Superpower

#121
post #29

Earlier quoted context omitted.

Regulatory capture is the OpenAI and Anthropic end goal, for certain. But I also think they exist in a sort of un-designed corporate narcissism, which is a common trait in bubble economies — I am not judging them particularly severely. Netscape under Clark and Andreessen and Sun under McNealy both fell into corporate narcissism: the belief that only they really mattered, that they were chosen, and that the world need…

Spot on. There's a certain level of drinking the kool-aid or getting high on their own supply. Anthropic is a lot worse than OpenAI but OpenAI had to go through rounds of shedding.

To be maximally fair to them, I think it is difficult to be one of the key businesses in a market bubble and not fall victim to this kind of thinking, especially when the continued inflation of the bubble depends on you — lots of people lose their shirts if you don't push hard to be "special".

But as you say, there is a measure of getting high on one's own supply now.

And there's the curious solipsistic energy of Sam Altman whimsically musing in public that it turns out his product is too expensive for people and they complain when you make the price realistic (when it possibly needs to be more expensive for OpenAI to survive).

They seem to believe that the ordinary rules either will not or somehow must not apply to them; it's increasingly bizarre to watch.

Maybe the people around pets.com were this bizarre; we didn't have so much livestreamed interview content to show us.

Re: Anthropic's Safety Superpower

#122
post #46
post #40

Earlier quoted context omitted.

I am just quoting the parent article. "What this degradation represented was both the capability and willingness of Anthropic to silently alter its models to achieve its policy preferences. In other words, Anthropic willfully validated some of its critics’ worst fears in terms of being a supply chain risk."

Again, hyperbole and assumption of evil intent because… they take precautions? Nice prose doesn’t dispense you from forming a sound hypothesis

The article makes a coherent argument:

a) Anthropic believe that AI is an extinction level risk and that they are the only leading AI lab which takes safety seriously. In combination this puts them in the position of believing that they are the only ones who can save the world, which is reasonable to call a god complex.

b) Anthropic are engaging in actions which aquire and consolidate power in the form of control over powerful AI.

c) "The history of brilliant people convinced they know what humanity needs is a sordid one, precisely because they have convinced themselves that their intentions are good, justifying actions that very much are not."

I'm describing claims from the article and would not word them so strongly myself. But this explicitly does not assume evil intent.

Re: Anthropic's Safety Superpower

#123

Earlier quoted context omitted.

For now I suspect however that the gigantic models are not needed and you will be able to do pretty much what you need in a specific domain with 120b or lower. There is so much trash in the frontier models. I don't need all the world's slam poetry for my coding tasks for example.

Wrong, mostly. Model capability is a function of model size. Raising the bar raises model performance in every domain. An "idiot savant" model that's overtrained for a specific domain would beat a generalist model of the same size. But scale the generalist up enough, and it'll trounce the specialist. Removing poetry data from a model training mix doesn't give you much - it might even cost you some performance - and "…

> Wrong, mostly.

> Model capability is a function of model size

Model effectiveness has improved across model sizes. You really should try the latest flash variants more. They have become my default for most tasks except for gnarly high-level planning.

Re: Anthropic's Safety Superpower

#124

> To that end, I can certainly buy the case that Fable/Mythos is in fact more capable when it comes to identifying and exploiting security issues This has been covered before: https://aisle.com/blog/ai-cybersecurity-after-mythos-the-jag... ( https://news.ycombinator.com/item?id=47732020 ) > Anthropic’s cautious roll-out was justified. The problem with publicly releasing models, however, is that guardrails can be jail…

We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights models. Is not We sent open weight models against a codebase to find vulnerabilities.

The second case is not what Anthropic did either, though. If you have their process internalized as "open freebsd, tell mythos 'find vulns', done" this is not what happened. They have a harness that went file-by-file, spawned a subagent for each file, told it to find vulns in that file, then a post-processing step (more on that in a sec).

In that sense: The AISLE replication still provides too much information to the model, but its not far off, and others have replicated Mythos' findings in a more clandestine manner on open source models. Some were totally capable of finding the same vulns Mythos found back in ~March (and today, the new Kimi K2.7 is looking extremely good, very little doubt it could do it).

The critical difference is that post-processing: the Mythos model/harness has some step to induce Mythos to actually exploit the vulnerability, leveraging its ability to do so as a ranking mechanism. Anthropic inferred that this led Mythos to discover vulnerabilities nothing else could discover, which is not true, and Anthropic should be held accountable for this weird artifact of that communication. However:

- An OSS model might find the vulnerability but rank it as a 3/10. Mythos finds it, chains it with a second vulnerability, now suddenly its an 8/10.

- An OSS model might find the vulnerability, alongside fifty other vulnerabilities. The operator ignores all of them.

The problem with automated vulnerability detection, including with LLMs, is that they find the haystack, not the needle. Every piece of hay might be a vulnerability, but whether its worthy of fixing is another matter. Mythos does represent a meaningful improvement; it better finds the needle.

Re: Anthropic's Safety Superpower

#125

Earlier quoted context omitted.

> The whole thesis falls apart though. You can't be on your way to "power over everything" and get distilled into free Chinese models within months. Pick one. But is that last part actually true though? Sure, there might be 600B+ models available for download and local inference if you have the hardware, but does the users who use Anthropic switch over to those even if they're available even as hosted models? Seems l…

> Anthropic and Claude remains very popular among the people who use LLMs Only because someone else is paying the bills. I use Claude Opus at work because my employer pays for the tokens and encourages me to do it. At home, I use DeepSeek Flash. It's not as good, but it's maybe 0.7 quality for 0.001 cost.

What's the speed on DeepSeek Flash? And what provider?

Re: Anthropic's Safety Superpower

#126
post #32

Earlier quoted context omitted.

That’s really what Dario wants. Let’s hope he doesn’t get it

what Dario wants is to retain any influence whatsover on how the research progresses before the inevitable nationalization of the frontier. he gets to keep the N-2 tech and maybe influence the N-1 tech, but the only influence on the frontier he has is today; whatever he imprints in the pipeline the government takes over. IOW I don't think he thinks in the same categories as most folks here.

> ...the research progresses before the inevitable nationalization of the frontier.

Hacker News has been telling me America beats China at "innovation" because of the "freedoms" - especially frew enterprise. I wonder how a nationalized frontier lab would perform.... Andhow the non-citizen researchers would feel about working for the US government that doesn't trust them to use frontier models.

Re: Anthropic's Safety Superpower

#127

> The entire Anthropic origin story is rooted in the founders’ belief that OpenAI wasn’t taking safety seriously enough; the company believes that only they can control AI, and that because they uniquely care about safety, they are justified in trying to control everyone else, up to and including the U.S. government. Anthropic believes they have the responsibility to guard their tools from mis-use. That is all. They…

[deleted]

Re: Anthropic's Safety Superpower

#128
post #10

Earlier quoted context omitted.

This is how the US gov does business now, capricious and vengeful. Textbook retaliation for not letting them use an abliterated version of Claude in weapons systems. This effectively renders any US closed model useless for any foreign company. Could happen to OpenAI, Google, etc. Too much of a risk to implement something that can be yanked out because the company didn’t behave the way they want. Looks like it’s time…

This is a suicide shot for the American economy. The numbers only lined up for AI to rescue the USA from its debt if it captured a significant portion of the world's AI spend, and while it was a longshot before, there's basically zero percent chance the world trusts American AI when the government is pulling strings.

> The numbers only lined up for AI to rescue the USA from its debt if it captured a significant portion of the world's AI spend

The numbers lined up if those companies created something resembling AGI, the USA companies captured a large share of the world, and there was lack of competition so those companies could capture a large share of the value.

None of those items were ever going go happen.

Re: Anthropic's Safety Superpower

#129
post #64

Perhaps they should consider leaving the US. Pretty clearly the descent into a corrupt autocracy is having real consequences.

Oh please, the earlier spat with the Trump admin was the best thing that ever happened to Anthropic. Before that, Claude was really only well-known in developer circles, not the wider normie-sphere. After Anthropic got the "Trump hates them, so it MUST be good!" stamp of approval, the company's recognition and popularity took off. This too, will end up being a good thing for them. The ban will end up getting lifted d…

If you’re implying that the government is in on it and is doing this stuff intentionally in order to boost Anthropic, that’s ridiculous.

Re: Anthropic's Safety Superpower

#130

Earlier quoted context omitted.

Wrong, mostly. Model capability is a function of model size. Raising the bar raises model performance in every domain. An "idiot savant" model that's overtrained for a specific domain would beat a generalist model of the same size. But scale the generalist up enough, and it'll trounce the specialist. Removing poetry data from a model training mix doesn't give you much - it might even cost you some performance - and "…

> Wrong, mostly. > Model capability is a function of model size Model effectiveness has improved across model sizes. You really should try the latest flash variants more. They have become my default for most tasks except for gnarly high-level planning.

"Capability per parameter" is rising, but parameter count remains an advantage. And small models remain bad, because "good" is a rapidly moving target.

A 2026 4B beats 2024 4B, but both are far behind the contemporary frontier. Which makes them bad. There is no such thing as "too much capability" - a "good" model is whatever the current frontier is.

In 2024, a "good" model is one that can be trusted to write a 800 line script. In 2026, it's a model that can be trusted to do gnarly high-level planning and execution both. In 2028, it's going to be something like a model you can point at an extremely involved task, abandon, and have it report back with a "done" in 3 weeks.

Post reply on HN