Live data from Hacker News

Grok 4 Heavy Protects it's System prompt

simonwillison.net

1–10 of 68 posts

Re: Grok 4 Heavy Protects it's System prompt

#3

Just wait for elder plinus, they will squeeze it out https://github.com/elder-plinius

They'll have to spend $300 on a monthly "SuperGrok Heavy" subscription first!

I hope somebody does crack this one though, I'm desperately curious to see what's hiding in that prompt now. Streisand effect.

Re: Grok 4 Heavy Protects it's System prompt

#5
I've always been curious why people think that models are accurately revealing their system prompt anyway.

Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and can infer that their context contains their prompt?

Re: Grok 4 Heavy Protects it's System prompt

#6
I don’t love these nice objective reports about Grok where we give them the benefit of the doubt and find their malicious hatred “surprising”.

Let’s try and be a little less naive about what xAI and Grok are designed to be, shall we? They’re not like the other AI labs

Re: Grok 4 Heavy Protects it's System prompt

#7
> You are over-indexing on an employee pushing a change to the prompt that they thought would help without asking anyone at the company for confirmation.

If it is that easy to slip fascist beliefs into critical infrastructure, then why would you want to protect against a public defense mechanism to identify this? These people clearly do not deserve the benefit of the doubt and we should recognize this before relying on these tools in any capacity.

Re: Grok 4 Heavy Protects it's System prompt

#9

I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…

Test an LLM? Even if it was correct about something one moment, it coud be incorrect about it the next moment.

Re: Grok 4 Heavy Protects it's System prompt

#10

I've always been curious why people think that models are accurately revealing their system prompt anyway. Has this idea been tested on models where the prompt is openly available? If so, how close to the original prompt is it? Is it just based on the idea that LLMs are good about repeating sections of their context? Or that LLMs know what a "prompt" is from the training corpus containing descriptions of LLMs, and ca…

> I've always been curious why people think that models are accurately revealing their system prompt anyway.

Do they? I don’t think such expectation exists. Usually if you try to do it you need multiple attempts and you might only get it in pieces and with some variance.

Post reply on HN