Live data from Hacker News

Claude 4 System Card

simonwillison.net

171–180 of 264 posts

Re: Claude 4 System Card

#171
post #137
post #86

> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.

I personally cancelled my Claude sub when they had an employee promoting this as a good thing on Twitter. I recognize that the actual risk here is probably quite low, but I don't trust a chat bot to make legal determinations and that employees are touting this as a good thing does not make me trust the company's judgment

>promoting this as a good thing

This is literally completely opposite of what happened. Then entire point is that this is bad, unwanted, behavior.

Additionally, it has already been demonstrated that every other frontier model can be made to behave the same way given the correct prompting.

I recommend the following article for an in depth discussion [0]

[0] https://thezvi.substack.com/p/claude-4-you-safety-and-alignm...

Re: Claude 4 System Card

#172
post #162
post #113

I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...

i found it oddly reassuring/decontextualizing to search/replace Claude with "your outie" + its nice to read in a markdowny format https://gist.github.com/swyxio/f207f99cf9e3de006440054563f6c...

lmao that's funny cause after seeing claude 4 code for you in zed editor while following it, it kinda feels like -the work is misteryous and interesting- level of work.

Re: Claude 4 System Card

#173

Earlier quoted context omitted.

Truly fascinating, thanks for this. What I find a little perplexing is when AI companies are annoyed that customers are typing "please" in their prompts as it supposedly costs a small fortune at scale yet they have system prompts that take 10 minutes for a human to read through.

Why not just strip “please” from the user input?

You can’t strip arbitrary words from the input because you can’t assume their context. The word could be an explicit part of the question or a piece of data the user is asking about.

Re: Claude 4 System Card

#174
post #113

I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...

I like reading the system prompt because I feel it would have to be human-written for sure, which is something I can never be sure of for all other text on the Internet. Or maybe not!

Re: Claude 4 System Card

#175
post #113

I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...

I like reading the system prompt because I feel it would have to be human-written for sure, which is something I can never be sure of for all other text on the Internet. Or maybe not!

I have absolutely iterated on system prompts with the help of LLMs before, so while system prompts generally will at the very least be heavily human curated, you can't assume that they are free of AI influence.

Re: Claude 4 System Card

#176

Earlier quoted context omitted.

Truly fascinating, thanks for this. What I find a little perplexing is when AI companies are annoyed that customers are typing "please" in their prompts as it supposedly costs a small fortune at scale yet they have system prompts that take 10 minutes for a human to read through.

Why not just strip “please” from the user input?

Why not just strip "" from the user input?

Re: Claude 4 System Card

#177

Earlier quoted context omitted.

It'd run in to all sorts of issues. Although AI companies losing money on user kindness is not our problem; it's theirs. The more they want to make these 'AIs' personable the more they'll get of it. I'm tired of the AIs saying 'SO sorry! I apologize, let me refactor that for you the proper way' -- no, you're not sorry. You aren't alive.

The obsequious default tone is annoying, but you can always prepend your requests with something like "You are a machine. You do not have emotions. You respond to exactly my questions, no fluff, just answers. Do not pretend to be a human."

I would also add "Be critical."

Re: Claude 4 System Card

#178
post #172
post #162

Earlier quoted context omitted.

i found it oddly reassuring/decontextualizing to search/replace Claude with "your outie" + its nice to read in a markdowny format https://gist.github.com/swyxio/f207f99cf9e3de006440054563f6c...

lmao that's funny cause after seeing claude 4 code for you in zed editor while following it, it kinda feels like -the work is misteryous and interesting- level of work.

Even the feeling of "this feels right" is there.

Oh no, are we the innies?

Re: Claude 4 System Card

#179
I read the whole thing, thank you for linking it Simon!

Notable to me is that Sonnet is really good at some coding use cases, better than Opus. It would make sense to me to distill Sonnet with an eye toward coding use cases - faster and cheaper - but I’m surprised it’s genuinely better, and it appears to be (slightly but measurably) better for some stuff.

Re: Claude 4 System Card

#180
post #84

Earlier quoted context omitted.

I think it was Larry Niven, quite a few decades ago, that had SF stories where AIs were only good for a few months before becoming suicidal...

Do you have any specific references? I’ve often wondered if human level intelligence might inevitably be plagued by human level neurosis and psychosis.

Sorry, fuzzy memory. I was going to write "six months", that's what stuck with me.

Not one of the mainline "Known Space" stories, if it was Niven at all. Maybe the suggestion about Frank Herbert in another comment is right, I also read a lot by him besides Dune - I particularly appreciated the Bureau of Sabotage concept ...

Post reply on HN