Grok 4 Heavy Protects it's System prompt
simonwillison.net
Grok 4 Heavy Protects it's System prompt
1–10 of 68 posts
Re: Grok 4 Heavy Protects it's System prompt
#2Re: Grok 4 Heavy Protects it's System prompt
#3Just wait for elder plinus, they will squeeze it out https://github.com/elder-plinius
I hope somebody does crack this one though, I'm desperately curious to see what's hiding in that prompt now. Streisand effect.
Re: Grok 4 Heavy Protects it's System prompt
#4Re: Grok 4 Heavy Protects it's System prompt
#5Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and can infer that their context contains their prompt?
Re: Grok 4 Heavy Protects it's System prompt
#6Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs
Re: Grok 4 Heavy Protects it's System prompt
#7If it is that easy to slip fascist beliefs into critical infrastructure, then why would you want to protect against a public defense mechanism to identify this? These people clearly do not deserve the benefit of the doubt and we should recognize this before relying on these tools in any capacity.
Re: Grok 4 Heavy Protects it's System prompt
#8Re: Grok 4 Heavy Protects it's System prompt
#9I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…
Re: Grok 4 Heavy Protects it's System prompt
#10I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…
Do they? I don’t think such expectation exists. Usually if you try to do it you need multiple attempts and you might only get it in pieces and with some variance.