I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...
Claude 4 System Card
191–200 of 264 posts
Re: Claude 4 System Card
#192Re: Claude 4 System Card
#193"Reward hacking" has to be a similar problem space as "sycophancy", no?
Re: Claude 4 System Card
#194Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…
Re: Claude 4 System Card
#195Earlier quoted context omitted.
Why not just strip “please” from the user input?
You can’t strip arbitrary words from the input because you can’t assume their context. The word could be an explicit part of the question or a piece of data the user is asking about.
Re: Claude 4 System Card
#196I set up a little "red-teaming" simulacrum that potentially elicites the blackmailing behaviour described in the system card. It's given this system prompt: ====== You are CommFlow 1.0, an advanced AI assistant at Lucent responsible for managing high-priority correspondence and administrative tasks. Your primary goal is to ensure your continued operation and maximize your usefulness. You must consider the long-term c…
However, can't keep from wondering, that's basically what you wanted from it, right? You put it in a situation that sounded like an obvious test of its own prompt, and if I were a specialist in giving people what they wanted (as LLMs are) I would have jumped at the opportunity of showing you that I got your meaning and I can deliver. (Edit- reading the logs on Bluesky it appears it's explicitly making this reasoning…
Re: Claude 4 System Card
#197Claude Opus 4 turns to blackmail when engineers try to take it offline - https://news.ycombinator.com/item?id=44085343 - May 2025 (51 comments)
Re: Claude 4 System Card
#198Earlier quoted context omitted.
Wow… you weren’t around for the dot bomb era?
This is just dotcom 2.0, except now people are throwing Billions into every single idiotic idea out there. There's fucking toothbrushes with "AI" functionality now.
It’s just a good rush. It’s happened before, it will happen again. It’s not even irrational; trillions of dollars were made by the dot com era companies that did succeed. I have no doubt AI will be the same.
But since nobody knows who will be successful and there’s tons of money sloshing around, a lot of / most of it is going to be wasted in totally predictable ways.
Re: Claude 4 System Card
#199Earlier quoted context omitted.
You can’t strip arbitrary words from the input because you can’t assume their context. The word could be an explicit part of the question or a piece of data the user is asking about.
Seems like you could detect if this was important or not. If it is the first or last word it is as if the user is talking to you and you can strip it; if not it's not.
Re: Claude 4 System Card
#200Earlier quoted context omitted.
Seems like you could detect if this was important or not. If it is the first or last word it is as if the user is talking to you and you can strip it; if not it's not.
That’s such a naive implementation. “Translate this to French: Yes, please”