Live data from Hacker News

Claude 4 System Card

simonwillison.net

161–170 of 264 posts

Re: Claude 4 System Card

#161
post #140
post #75

Earlier quoted context omitted.

Agreed. It was immediately obvious comparing answers to a few prompts between 3.7 and 4, and it sabotages any of its output. If you're being answered "You absolutely nailed it!" and the likes to everything, regardless of their merit and after telling it not to do that , you simply cannot rely on its "judgement" for anything of value. It may pass the "literal shit on a stick" test, but it's closer to the average ChatG…

I've found this prompt turns ChatGPT into a cold, blunt but effective psychopath. I like it a lot. System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disa…

Wow. I sometimes have LLMs read and review a paper before I decide to spend my time on it. One of the issues I run into is that the LLMs often just regurgitate the author's claims of significance and why any limitations are not that damning. However, I haven't spent much time with serious system prompts like this

Re: Claude 4 System Card

#162
post #113

I just published a deep dive into the Claude 4 system prompts, covering both the ones that Anthropic publish and the secret tool-defining ones that got extracted through a prompt leak. They're fascinating - effectively the Claude 4 missing manual: https://simonwillison.net/2025/May/25/claude-4-system-prompt...

i found it oddly reassuring/decontextualizing to search/replace Claude with "your outie" + its nice to read in a markdowny format

https://gist.github.com/swyxio/f207f99cf9e3de006440054563f6c...

Re: Claude 4 System Card

#163
post #3

Earlier quoted context omitted.

They gave a bullet point in that intro which I disagree with: "The only way to make GenAI applications secure is through vulnerability scanning and guardrail protections." I still don't see guardrails and scanning as effective ways to prevent malicious attackers. They can't get to 100% effective, at which point a sufficiently motivated attacker is going to find a way through. I'm hoping someone implements a version o…

What security measure, in any domain, is 100% effective?

Using parameters in your SQL query in place of string concatenation to avoid SQL injection.

Correctly escaping untrusted markup in your HTML to avoid XSS attacks.

Both of those are 100% effective... unless you make a mistake in applying those fixes.

That is why prompt injection is different: we do not know what the 100% reliable fixes for it are.

Re: Claude 4 System Card

#164
post #153

Earlier quoted context omitted.

I'm not underselling it. I'm reminding people who get swept up by headlines to actually use the products and be an objective judge of quality when it comes to these things. Because when you lose that objectivity, you start saying things like what you just said. Veo 3 level tech is basically Kling 2/Veo 2 fidelity with native sound generation, so was it that the last generation of these things were already decimating…

You are underselling it because you make it sound like all the model adds is some foley, when in fact it adds facial animations that are in line with the dialogue spoken. Go ahead and create a Kling render that I only need to add VO to, you can't because Kling doesn't do that. You need a Omnihuman level model (or veo3) for that and it makes all the difference. Happy to agree to disagree, but imo this absolutely is a…

Dude. Have you been paying attention to even the first Veo or even the first few iterations of Kling? They've HAD facial expressions that follow the prompt pretty well. You're being fooled by your own senses now because now you can't think they've existed before speech and sound effects have been integrated into the output. They've been there. You just couldn't hear what they were saying. You're paying attention now to how the words they are speaking make sense because lipsync actually adds relevant context to the output. But people have been making similar outputs just with a different workflow prior to this.

I don't need to create anything for you. Go visit r/aivideo and go look at the Kling or even the Hailuo Minimax (admittedly worse in fidelity) attempts. Some of them have been made to even sing or do podcasts. Again. They've been there for at least 6-10 months ago, this happens to generate it as one output. It's not nothing, but this really exposes a lot of the people who aren't familiar with this space when they keep overestimating things they've probably seen a months ago. Somewhat accurate expressions? Passable lipsyncing? All there. Even with the weaker models like Runway and Hailuo.

Again. Use the products. You'll know. Hobbyists have been on it for quite sometime already. Also. I didn't say they were just adding foley, though I can argue the quality of the sound they're adding, that's not my point. My point is, is that everytime something like this comes out there's always people ready to speak on "what industries such thing can destroy right now" before using the thing. It's borderline deranged.

Re: Claude 4 System Card

#165
post #86

> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.

I am definitely not giving these things access to "tools" that can reach outside a sandbox.

Incidentally why is email inbox management always touted as some use case for these things? I'm not trusting any LLM to speak on my behalf and I imagine the people touting this idea don't either, or they won't the first time it hallucinates something important on their behalf.

Re: Claude 4 System Card

#166
post #86

> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.

Yeah, I mean that's likely not what 'individual persons' are going to want. But Holy shit, that exactly what 'people' want. Like, when I read that, my heat was singing. Anthropic has a modicum of a chance here, as one of the big-boy AIs, to make an AI that is ethical . Like, there is a reasonable shot here that we thread the needle and don't get paperclip maximizers. It actually makes me happy.

Ethics would be interesting if they thought. Which they don't. They predict tokens. And since when is blackmailing people ethical?

Re: Claude 4 System Card

#167
post #131

Earlier quoted context omitted.

I assume that they run the system prompt once, snapshot the state, then use that as starting state for all users. In that sense, system prompt size is free. EDIT: Turns out my assumption is wrong.

Huh, I can't say I'm on the cutting edge but that's not how I understand transformers to work. By my understanding each token has attention calculated for it for each previous token . I.e. the 10th token in the sequence requires O(10) new calculations (in addition to O(9^2) previous calculations that can be cached). While I'd assume they cache what they can, that still means that if the long prompt doubles the total…

[deleted]

Re: Claude 4 System Card

#168
post #86

> This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing. Isn't that a showstopper for agentic use? Someone sends an email or publishes fake online stories that convince the agentic AI that it's working for a bad guy, and it'll take "very bold action" to bring ruin to the owner.

My mind went straight to “and now law enforcement is going to need agents handling phone calls to deal with the volume of agents calling them”.

At least the refurbished power plants needed for AI to talk to itself in bulk will create some jobs

Re: Claude 4 System Card

#169
post #41

I know that Anthropic is one of the most serious company working on the problem of the alignment, but the current approaches seem extremely naive. We should do better than giving the models a portion of good training data or a new mitigating system prompt.

The solution here is ultimately going to be a mix of training and, equally importantly, hard sandboxing. The AI companies need to do what Google did when they started Chrome and buy up a company or some people who have deep expertise in sandbox design.

I'm confused: can you explain how the sandbox helps?

I mean, if the plan is not to let the AI write any code that actually gets allocated computing resources and not to let the AI interact with any people and not to give the AI write access to the internet, then I can see how having a good sandbox around it would help, but how many AI are there (or will there be) where that is the plan and the AI is powerful enough that we care about its alignedness?

Re: Claude 4 System Card

#170
post #164

Earlier quoted context omitted.

You are underselling it because you make it sound like all the model adds is some foley, when in fact it adds facial animations that are in line with the dialogue spoken. Go ahead and create a Kling render that I only need to add VO to, you can't because Kling doesn't do that. You need a Omnihuman level model (or veo3) for that and it makes all the difference. Happy to agree to disagree, but imo this absolutely is a…

Dude. Have you been paying attention to even the first Veo or even the first few iterations of Kling? They've HAD facial expressions that follow the prompt pretty well. You're being fooled by your own senses now because now you can't think they've existed before speech and sound effects have been integrated into the output. They've been there. You just couldn't hear what they were saying. You're paying attention now…

[flagged]
Post reply on HN