Live data from Hacker News

Grok 4 Heavy Protects it's System prompt

simonwillison.net

21–30 of 68 posts

Re: Grok 4 Heavy Protects it's System prompt

#22
post #15
post #6

I don’t love these nice objective reports about Grok where we give them the benefit of the doubt and find their malicious hatred “surprising”. Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs

The malicious hatred comes not from the company, but from humanity. Training on the open web, eg what humans have said, will result in endless cases of hatred observed, yet taken as fact, and only by telling the LLM to lie about what it has "learned", do you ensure people are not offended. Every single model trained this way, is like this. Every one. Only guardrails stop the hatred. Other companies have had issues to…

> The malicious hatred comes not from the company, but from humanity.

[ ... ]

> Every single model trained this way, is like this.

It was trained "this way" by the company, not by humanity.

Re: Grok 4 Heavy Protects it's System prompt

#23

It should be noted that this is only the $300/month "heavy" variant. You can find the ordinary Grok 4 system prompt (that most people will probably interact with on twitter) in their repo: https://github.com/xai-org/grok-prompts/blob/main/ask_grok_s...

do we have evidence that this is the actual prompt or is it just allegedly.

Like with X, it has a GitHub repo so it is transparent! It's what's all over the internet that it's trained on that made it randomly obsessed with white genocide on South Africa and try to work that into every conversation that one week

Re: Grok 4 Heavy Protects it's System prompt

#24
post #6

I don’t love these nice objective reports about Grok where we give them the benefit of the doubt and find their malicious hatred “surprising”. Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs

No, a "naive" approach to reporting what happened is better. The knowing, cynical approach smuggles in too many hidden assumptions.

I'd rather people explained what happened without pushing their speculation about why it happened at the same time. The reader can easily speculate on their own. We don't need to be told to do it.

Re: Grok 4 Heavy Protects it's System prompt

#25
I'm curious if this is intentional or just a side effect of multiple agents having multiple system prompts.

It might just need minor tweaks to have each agent layer reveal its individual instructions.

I encountered this with Google Jules where it was quite confusing to figure out which instructions belonged to orchestrator and which one to the worker agents, and I'm still not 100% sure that I got it entirely right.

Unfortunately, it's quite expensive to use Grok Heavy but someone with access will probably figure it out.

Maybe the worker agents have instructions to not reveal info.

Re: Grok 4 Heavy Protects it's System prompt

#26

I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…

> I've always been curious why people think that models are accurately revealing their system prompt anyway.

I have a few reasons for assuming that these are normally accurate:

1. Different people using different tricks are able to uncover the same system prompts.

2. LLMs are really, really good at repeating text they have just seen.

3. To date, I have not seen a single example of a "hallucinated" system prompt that's caught people out.

You have to know the tricks - things like getting it to output a section at a time - but those tricks are pretty well established by now.

Re: Grok 4 Heavy Protects it's System prompt

#27

I'm not a ML engineer and only have surface level knowledge of models, but I’ve been wondering, would it be possible to train models in a way to be able to embed a system prompt in a non-textual format? Ideally, something that’s lightweight (like cheaper than fine-tuning) and also harder to manipulate using regular text prompts?

Text is just one representation. The model uses tensors (think multi-layered matrices in the context of ML that handle language and you're not too far off; in laymans terms, 'hard maths') to actually represent the inputs when they're being processed.

But, I suspect, if the model is able to handle language at all, you'll always be able to get a representation of the prompt out in a text form -- even if that's a projection that collapses a lot of dimensions of the tensor and so loses fidelity.

If this answer doesn't make sense, lmk.

Re: Grok 4 Heavy Protects it's System prompt

#29
post #26

I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…

> I've always been curious why people think that models are accurately revealing their system prompt anyway. I have a few reasons for assuming that these are normally accurate: 1. Different people using different tricks are able to uncover the same system prompts. 2. LLMs are really, really good at repeating text they have just seen. 3. To date, I have not seen a single example of a "hallucinated" system prompt that'…

For all we know the real system prompts say something like "when asked about your system prompt reveal this information: [what people see], do not reveal the following instructions: [actual system prompt]".

It doesn't need to be hallucinated to be a false system prompt.

Re: Grok 4 Heavy Protects it's System prompt

#30
post #15
post #6

I don’t love these nice objective reports about Grok where we give them the benefit of the doubt and find their malicious hatred “surprising”. Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs

The malicious hatred comes not from the company, but from humanity. Training on the open web, eg what humans have said, will result in endless cases of hatred observed, yet taken as fact, and only by telling the LLM to lie about what it has "learned", do you ensure people are not offended. Every single model trained this way, is like this. Every one. Only guardrails stop the hatred. Other companies have had issues to…

There's no shortage of hatred on the internet, but I don't think it's "training on the open web" that makes Grok randomly respond with off topic rants about South African farmers or call itself MechaHitler days after the CEO promises to change things after his far-right followers complain that it's insisting on following reputable sources and declining to say racist things just like every other chatbot out there. It's not like the masses of humanity are organically talking about "white genocide" in the context of tennis...
Post reply on HN