Live data from Hacker News

Claude 4 System Card

simonwillison.net

261–264 of 264 posts

Re: Claude 4 System Card

#261

Earlier quoted context omitted.

I'm confused: can you explain how the sandbox helps? I mean, if the plan is not to let the AI write any code that actually gets allocated computing resources and not to let the AI interact with any people and not to give the AI write access to the internet, then I can see how having a good sandbox around it would help, but how many AI are there (or will there be) where that is the plan and the AI is powerful enough t…

The problems here aren't different to restricting malicious or hacked employees, or malicious or hacked third party libraries. You start with the low hanging fruit: run tool commands inside a kernel sandbox that switches off internet access and then re-provide access only via an HTTP proxy that implements some security policies. For example, instead of providing direct access to API keys you can give the AI a fake on…

This conversation began as a conversation about Claude, which has access to 100s of 1000s of people with no training and no interest in learning about how to prevent Claude from doing damage to society. That makes it materially different from a library because even if an intruder can subvert a library running on servers serving 100s of 1000s of users, e.g., a library for compressing files is very unlikely to be able to start having conversations with a large fraction of those users without someone noticing that something is very wrong.

Although I concede that there are some applications of AI that can be made significantly safer using the measures you describe, you have to admit that those applications are fairly rare and emphatically do not include Claude and its competitors. For example, Claude has plentiful access to computing resources because people routinely ask it to write code, most of which will go on to be run (and Claude knows that). Surely you will concede that Anthropic is not about to start insisting on the use of a sandbox around any code that Claude writes for any paying customer.

When Claude and its competitors were introduced, a model would reply to a prompt, then about a second later it lost all memory of that prompt and its reply. Such an LLM of course is no great threat to society because it cannot pursue an agenda over time, but of course the labs are working hard to create models that are "more agentic". I worry about what happens when the labs succeed at this (publicly stated) goal.

Re: Claude 4 System Card

#263
post #72

Earlier quoted context omitted.

the big difference is the capability to think during tool calls. this is what makes openAI o3 lookin like magic

Yeah, I've noticed this with Qwen3, too. If I rig up a nonstandard harness than allows it to think before tool calls, even 30B A3B is capable of doing low-budget imitations of the things o3 and similar frontier models do. It can, for example, make a surprising decent "web research agent" with some scaffolding and specialized prompts for different tasks. We need to start moving away from Chat Completions-style tool ca…

> If I rig up a nonstandard harness than allows it to think before tool calls

What does that require? (I'm extremely, extremely new to all this.)

Re: Claude 4 System Card

#264
post #236

Earlier quoted context omitted.

I didn't read it as attempted comedy. I am genuinely dismayed by how easy it is for grifters to continue to find victims long after being exposed.

"Attempted comedy" is the most charitable take I could give it (and truthfully, I implied something else). My point is, their statement is quite obviously wrong, but it sure sounds nice . If you don't agree, I challenge you to provide that track record "of non-stop lying, scamming and bllsh*tting right in people's faces for years". Like, for real. I'm not defending 'sama here; I'm not a fan of his either (but neither…

Name an example of something impressive HE built or did or said that was not a lie or scam?

You claim I'm "obviously" wrong. So where are the arguments?

Post reply on HN