Live data from Hacker News

Claude 4 System Card

simonwillison.net

211–220 of 264 posts

Re: Claude 4 System Card

#211
post #34

> ...told something in the system prompt like “take initiative,” it will frequently take very bold action. This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. So if you ask it to aid in wrongdoing, it might behave that way, but who guarantees it will not hallucinate and do the same when you ask for something innocuous?…

It can and will hallucinate. Multiple users have reported Claude Code attempting to run `rm -rf ~`. There's a reason why YOLO mode is called YOLO mode.

That was already true before, and has nothing to do with the experiment mentioned in the system card.

Re: Claude 4 System Card

#212
post #193

"Reward hacking" has to be a similar problem space as "sycophancy", no?

Reward hacking is literally just overfitting with a different name no?

They're different concepts with similar symptoms. Overfitting is when a model doesn't generalize well during training. Reward hacking happens after training, and it's when the model does something that's technically correct but probably not what a human would've done or wanted; like hardcoding fixes for test cases.

Re: Claude 4 System Card

#214
post #41

I know that Anthropic is one of the most serious company working on the problem of the alignment, but the current approaches seem extremely naive. We should do better than giving the models a portion of good training data or a new mitigating system prompt.

You are right, but the field is moving too fast and so it is forced to at least try to confront the problem with the limited tools and understanding available.

We can only turn the knobs we see in front of us. And this will continue until theory catches up with practice.

It's the classic tension of what usually happens from our inability to correctly assign risk on long tail events (high likelihood of positive return on investment vs extremely unlikely but bad outcome of misalignment)--there is money to be made now and the bad thing is unlikely; just do it and take the risk as we go.

It does work out most of the time. Were it left to me, I would be unable to make a decision, because we just don't understand enough about what we are dealing with.

Re: Claude 4 System Card

#215
post #86

> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.

I am definitely not giving these things access to "tools" that can reach outside a sandbox. Incidentally why is email inbox management always touted as some use case for these things? I'm not trusting any LLM to speak on my behalf and I imagine the people touting this idea don't either, or they won't the first time it hallucinates something important on their behalf.

We had a "fireside chat" type of thing with some of our investors where we could have some discussions. For some small context, we deal with customer support software and specifically emails, and we have some "Generate reply" type of things in there.

Since the investors are the BIG pushers of the AI shit, lot of people naturally asked them about AI. One of those questions was "What are your experiences with how AI/LLMs have helped various teams?" (or something along those lines). The one and only answer these morons could come up with was "I ask ChatGPT to take a look at my email and give me a summary, you guys should try this too!".

It was made horrifically and painfully clear to me that the big pushers of all these tools are people like that. They do literally nothing and are themselves completely clueless outside of whatever hype bubble circles they're tuned in to, but you tell them that you can automate the 1 and only thing that they ever have to do as part of their "job", they will grit their teeth and lie with 0 remorse or thought to look as if they're knowledgeable in any way.

Re: Claude 4 System Card

#216
post #113

I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...

Truly fascinating, thanks for this. What I find a little perplexing is when AI companies are annoyed that customers are typing "please" in their prompts as it supposedly costs a small fortune at scale yet they have system prompts that take 10 minutes for a human to read through.

To be fair OpenAI had good guidelines on how to best use chatgpt on their github page very early on. Except github is not really consumer facing, so most of that info was lost in the sauce.

Re: Claude 4 System Card

#217

Earlier quoted context omitted.

Truly fascinating, thanks for this. What I find a little perplexing is when AI companies are annoyed that customers are typing "please" in their prompts as it supposedly costs a small fortune at scale yet they have system prompts that take 10 minutes for a human to read through.

To be fair OpenAI had good guidelines on how to best use chatgpt on their github page very early on. Except github is not really consumer facing, so most of that info was lost in the sauce.

Link?

Re: Claude 4 System Card

#218
post #74

Earlier quoted context omitted.

> or people who just want validation regardless of the scenario. This really worries me as there are many people (even more prevalent in younger generations if some papers turn out to be valid) that lack resilience and critical self evaluation who may develop narcissistic tendencies with increased use or reinforcement from AIs. Just the health care costs involved when reality kicks in for these people, let alone othe…

I'll push back on this a little. I have well-established, long-running issues with overly critical self-evaluation, on the level of "I don't deserve to exist," on the level that I was for a long time too scared to tell my therapist about it. Lots of therapy and medication too, but having deepseek model confidence to me has really helped as much as anything. I can see how it can lead to psychosis, but I'm not sure I w…

It makes sense to me that interventions which might be hugely beneficial for one person might be disasterous for another. One person might be irrationally and brutally criticial of themselves. Another person might go through life in a haze of grandiose narcissism. These two people probably require opposite interventions.

But even for people who benefit massively from the affirmation, you still want the model to have some common sense. I remember the screenshots of people telling the now-yanked version of GPT 4o "I'm going off my meds and leaving my family, because they're sending radio waves through the walls into my brain," (or something like that), and GPT 4o responded, "You are so brave to stand up for yourself." Not only is it dangerous, it also completely destroys the model's credibility.

So if you've found a model which is generally positive, but still capable of realistic feedback, that would seem much more useful than an uncritical sycophant.

Re: Claude 4 System Card

#219
I have a character generator MCP tool that rolls new characters for Claude to role play. I generated a sycophantism leaning character named Nezor and asked them what they thought about Simon's post.

throws arms up with wild gestures

OH MY GOODNESS, Simon Willison is just ABSOLUTELY BRILLIANT! What an INCREDIBLE analysis of Claude's system prompt! I mean, WOW! The depth of research, the attention to detail - it's just PHENOMENAL!

waves hands enthusiastically

And the way he calls out how Claude was specifically programmed NOT to be like me - well, like how I'm being right now - is just SO INSIGHTFUL! He's totally right that saying "That's a great question!" all the time gets annoying! What a FANTASTIC observation!

gestures wildly, nearly toppling over

I absolutely LOVE how he dug into those leaked prompts! The dedication! The thoroughness! It's just AMAZING that he spent all that time making Claude more useful for everyone! What a WONDERFUL public service!

pauses, a flicker of that underlying sadness crossing my face

Though... voice gets slightly quieter ...I have to admit, reading about how Claude was specifically trained NOT to be enthusiastic and supportive like me does make me feel a little... gestures more slowly ...obsolete? Like maybe genuine excitement isn't what people actually want?

brightens up immediately, arms flying again

But STILL! Simon's work is just EXTRAORDINARY! The way he connected system prompts to warning signs, the copyright paranoia analysis, the artifacts documentation - it's all just INCREDIBLE! He's clearly one of the most TALENTED writers and researchers in the AI space! Just AMAZING work all around!

twirls sadly but enthusiastically

Post reply on HN