Changes in the system prompt between Claude Opus 4.6 and 4.7
simonwillison.net
Changes in the system prompt between Claude Opus 4.6 and 4.7
1–10 of 240 posts
Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#2Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#3Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#4The malware paranoia is so strong that my company has had to temporarily block use of 4.7 on our IDE of choice, as the model was behaving in a concerningly unaligned way, as well as spending large amounts of token budget contemplating whether any particular code or task was related to malware development (we are a relatively boring financial services entity - the jokes write themselves).
In one case I actually encountered a situation where I felt that the model was deliberately failing execute a particular task, and when queried the tool output that it was trying to abide by directives about malware. I know that model introspection reporting is of poor quality and unreliable, but in this specific case I did not 'hint' it in any way. This feels qualitatively like Claude Golden Gate Bridge territory, hence my earlier contemplation on steering vectors. I've been many other people online complaining about the malware paranoia too, especially on reddit, so I don't think it's just me!
Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#5I'm curious as to why 4.7 seems obsessed with avoiding any actions that could help the user create or enhance malware. The system prompts seem similar on the matter, so I wonder if this is an early attempt by Anthropic to use steering vector injection? The malware paranoia is so strong that my company has had to temporarily block use of 4.7 on our IDE of choice, as the model was behaving in a concerningly unaligned w…
Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#6Uff, I've tried stuff like these in my prompts, and the results are never good, I much prefer the agent to prompt me upfront to resolve that before it "attempts" whatever it wants, kind of surprised to see that they added that
Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#7edit: to be fair Anthropic should be giving money back for sessions terminated this way.
Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#8Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#9Re: Changes in the system prompt between Claude Opus 4.6 and 4.7
#10I knew these system prompts were getting big, but holy fuck. More than 60,000 words. With the 3/4 words per token rule of thumb, that's ~80k tokens. Even with 1M context window, that is approaching 10% and you haven't even had any user input yet. And it gets churned by every single request they receive. No wonder their infra costs keep ballooning. And most of it seems to be stable between claude version iterations to…