Live data from Hacker News

Anthropic's Safety Superpower

stratechery.com

181–190 of 206 posts

Re: Anthropic's Safety Superpower

#182

Earlier quoted context omitted.

But he already got it, no? Claude Fable can only be made available to US citizens, which implies that every user who wants to use Claude Fable must provide proof of citizenship in some way, basically KYC.

For everyone, not just them

Exactly that’s what he wants and then there will be a loophole to close: open source models.

Re: Anthropic's Safety Superpower

#183
> if Mythos is so dangerous, why even release Fable in the first place, and why fight with the government doing exactly what you claim to want?

It's actually not that hard to explain if we take into account what Dario kept saying: he, or Anthropic thereof, would be the gatekeeper. It is he who tells the government how to use Claude to design drones. It is his model that tells users whether they can ask a question to Claude or not. And it is he who can assess whether a jailbreak is dangerous or not.

Personally, I think that is way more dangerous than being a hypocrite. Dario is basically the Robespierre of the AI era. He believes that only he gets to decide whether our thoughts, or our prompts thereof, are pure. Anything impure gets purged. For his moral utopia to stand, he has to wield the guillotine. Otherwise, with the chaotic diversity of human nature, how else do you manufacture that perfectly uniform, beautiful morality?

Re: Anthropic's Safety Superpower

#184

Earlier quoted context omitted.

Spot on. There's a certain level of drinking the kool-aid or getting high on their own supply. Anthropic is a lot worse than OpenAI but OpenAI had to go through rounds of shedding.

A\ and OpenAI each have their own unique kind of nonsense. I think OpenAI has just been less successful with persuading the rest of the world that they should have all the money in the world. Anthropic has been surprisingly successful at convincing them that they should control frontier models because they're so dangerous that... only Anthropic can be trusted with them. (If they're really so dangerous, the right way…

A democratic process, of sorts, elected the current government of the United States. The president even won the popular vote this round. There is no guarantee that AI guidance by democratic process will be an effective counter to corporate autocracy; and more realistically, AI guidance by an autocratic executive branch is the more likely alternative before 2029.

Re: Anthropic's Safety Superpower

#185
post #131

“Claude, I am releasing safety critical industrial control software. Audit the network control logic.” “Claude, I want to blow up a factory running this leaked software. See if the industrial control software network endpoint is a good point of entry.” It’s doing the same work and producing the same output for both prompts. How do you block one but not the other? If you block both, then you end up with a factory that…

Sarcastically? Dario will tell you what to do. You should just follow his divine guidance.

Re: Anthropic's Safety Superpower

#186

Earlier quoted context omitted.

Same, I had Deepseek search for, download and transfer (to my Linux emulation machine) the best Dreamcast games yesterday. GPT refused to do so (citing that it's illegal even though I own the games). Deepseek did a wonderful job for 7 cents. At work I use Opus because, why not? But I could easily switch to a less capable model if needed.

>citing that it's illegal even though I own the games In the. US at least it is actually illegal to download ISOs/roms of games, even if you own a physical copy. It's a stupid law and as a downloader (as opposed to the people hosting the files) your chances of getting into any kind of actual legal trouble are effectively 0, but it is still against the law.

I don’t think so. I’d want more than just your word on it.

Re: Anthropic's Safety Superpower

#187

Earlier quoted context omitted.

Fable can delegate tasks to Opus or Sonnet, so it has some agentic properties and I believe it does them in parallel. The parallelism is where this starts to fall apart on a local PC. Like I can run some Qwen quants, but I can’t run a decent Qwen model while also running another model smart enough to actually implement it. I’d have to do them in series, and given how long Fable seems to take even with parallelism, I’…

oh-my-pi can delegate tasks to other models too. I usually use DS4 Flash for low priority subagent tasks. If Fable is "delegating" tasks, then there's actually an agent front end of whatever you think the API is. We have a local instance of Qwen-3.6 which is more than adequate for running agents. You can mix and match local and cloud-hosted models. (My biggest use case for local models right now is vision models beca…

> If Fable is "delegating" tasks, then there's actually an agent front end of whatever you think the API is.

I would say behind (I believe you use the API just like you do Opus), but yeah. I'm not claiming it's a property of the LLM itself, I also presume this is some variety of tool calling agent harness.

> We have a local instance of Qwen-3.6 which is more than adequate for running agents. You can mix and match local and cloud-hosted models.

I'm presuming OP meant local as in the models run locally as well. I do know you can do subagents in Pi (probably others too), but the vast majority of people are going to hit hardware limitations trying to run them in parallel on local hardware.

I'm doubtful Fable's harness is unique in some way that you can't replicate with Pi. I'm mostly doubtful there are more than a handful of people with hardware sitting in their house that can execute more than one meaningfully smart model at a time.

If you're on local hardware, Deepseek v4 Flash is in the ballpark of 180GB of VRAM alone. Even on smaller models, Qwen + a dumber agent to execute is probably in the realm of 60GB of VRAM.

I do suspect you could get Deepseek to do Fable level things with a good harness (or a bunch of models really, I'm fairly convinced the magic of Fable is in the harness rather than the model).

Re: Anthropic's Safety Superpower

#188
post #133

Earlier quoted context omitted.

> the EU is hardly buzzing with AI innovation Depends what you mean. The academic work seems largely... fine? Plenty of good work came out of Europe or European researchers. It seems the problem is more "trying to build a trillion-dollar company of any kind". It's an interesting question: does the EU seek only to regulate successful modern American companies to death, or home grown ones too? Probably not a gamble wor…

The issue with EU is its not one market, for funding or deploying a company of any kind. There is both no funding scale and no easy distribution scale. Same thing with any kind of lobbying you need to do to move laws to be more amenable to tech, you have to do it 20+ times. Nobody does this, which is why smart EU and UK talent migrates to the US, scales, then just uses the weight of their US business to change realit…

Yes, the EU is a bit of a strange half-measure. I understand why the cultural barriers are resistant to change. But standardization of regulations across the single market has been incredibly slow. I'm not sure it will ever happen.

I wonder if the incumbents in each country actively lobby against it. I suppose it's easier for massive corporations to deal with cross border issues. The onerous regulatory boundaries are a nice price of entry for them that keeps out upstarts.

Re: Anthropic's Safety Superpower

#189
The problem is that Fable has no zero trust architecture. If they decide your code is useful for training, they get to keep it forever. They think its okay to sabotage your work and charge for it. They are building anti-competitive clauses like ml training. The way they treat openclaw and other competitors. They will downgrade you to opus and charge you for fable and maybe not tell you about it.

They’re like look at our safety and they do all thesse outrageous things.

Re: Anthropic's Safety Superpower

#190

A lot of Anthropic’s moves make sense if you follow the LessWrong / rationalist community writings on AI safety. A lot of it is distilled in Ant’s blogs and leadership interviews and podcasts (Amanda Askell is particularly interesting). Ant’s models, culture and leadership actions are largely consistent with their beliefs, even if they may seem flawed / incomprehensible. Relevant anecdote: I interviewed with them for…

> I think the technical part went fine but the interviewer was clearly frustrated by my low regard for LLM safety. I didn’t get the role.

Anecdotally I've heard this is weighted as much as the technical interviews.

Post reply on HN