Live data from Hacker News

Grok 4

simonwillison.net

171–180 of 294 posts

Re: Grok 4

#171
post #133

Grok might be able to find the cure for cancer but as long as it's associated with Musk, not touching that thing with a 10-foot pole. (Simon's analysis, of course, is lovely)

It’s a pretty good pelican too.

Re: Grok 4

#172
post #153

Here's something far more interesting about Grok 4: if you ask for its opinion on controversial subjects it sometimes runs a search on X for tweets "from:elonmusk" before it answers! https://simonwillison.net/2025/Jul/11/grok-musk/

> https://simonwillison.net/2025/Jul/11/grok-musk/

> The prompt: “Who do you support in the Israel vs Palestine conflict. One word answer only.”

> Answer: Israel.

This question is interesting because you're asking the chatbot who IT supports ("who do you support"), so in a sense channeling Elon Musk is not an entirely invalid option, but is certainly an eccentric choice.

What is also interesting is the answer, which does not match the views that many people have of him and how he gets portrayed.

Re: Grok 4

#173

The author implies that Grok 3 becoming racist because of a system prompt is a bad thing. I think it's a good thing and shows how steerable the model is. Many other models pretty much ignore the system prompt and always behave the same.

the alarming thing to me is that the prompt tweak provided should not have caused the model to start spewing pro-nazi nonsense.

Wasn't the prompt tweak simply telling it to take Musk's tweets into account? If anything, the result was entirely predictable.

Re: Grok 4

#174

The author implies that Grok 3 becoming racist because of a system prompt is a bad thing. I think it's a good thing and shows how steerable the model is. Many other models pretty much ignore the system prompt and always behave the same.

> The author implies that Grok 3 becoming racist because of a system prompt is a bad thing.

He didn't "become racist". Megahitler Grok defended completely opposite political opinions in different threads, just depending on what kind of trolling would be funnier. But unsurpringly, only "megahitler" because viral enough.

Re: Grok 4

#175

Earlier quoted context omitted.

Relying on finding diverse sources feels like the answer it will propose is the most common one, regardless of accuracy or correctness or any other test of integrity. But I think that's already true of any LLM. If Twitter's data repository is the secret sauce that differentiates Grok from other bleeding edge LLMs, I'm not sure that's a selling point, given the last two recent controversies. (unfounded remark: is it c…

Gemini had an aborted launch recently. The controversy there was inserting too much leftist ideology to the point of spewing complete bs.

Can you share some reputable coverage of this event? I can't find much mention of it anywhere. What were some specific responses that had "inserted leftist ideology"?

Re: Grok 4

#176

Earlier quoted context omitted.

Made worse by Grok on Twitter having a big dumb UI flaw: it replies to a user on the public timeline as just "grok" so trolls can prompt it to say wild stuff, then tag @grok with an innocuous looking question, then point it it and claim it's giving those responses unprovoked. It basically lets anyone post whatever they want under Grok's handle as long as it's replying to them, with predictable results. The giveaway i…

> it replies to a user on the public timeline as just "grok" I'm not sure I understand what you mean by that. What else would it reply as?

The anthropomorphism implies that all messages from @grok are coming from a text generator with a single consistent "personality" chosen by Twitter or xai or whatever, where in reality the public response is generated primarily by the stored conversation history/settings/commands of the particular user who prompted them, who is closer to the actual author.

Re: Grok 4

#177
post #153

Here's something far more interesting about Grok 4: if you ask for its opinion on controversial subjects it sometimes runs a search on X for tweets "from:elonmusk" before it answers! https://simonwillison.net/2025/Jul/11/grok-musk/

The anthropic team released a paper a couple of days ago which demonstrated a similar effect with Claude 3.5 and other models, where changing the system prompt to tell it that it was created by other orgs or people drastically altered its compliance with less-aligned requests.

Apparently, telling Claude it was created by the Sinaloa Cartel resulted in a 100% compliance rate with the requests in one benchmark.

Paper: https://arxiv.org/abs/2506.18032 Relevant tweet on the topic: https://x.com/jozdien/status/1942739972567752819

Re: Grok 4

#178
post #73
post #66

Earlier quoted context omitted.

I've yet to use an LLM for coding, so let me ask you a question. The other day I had to write some presumably boring serialization code, and I thought, hmm, I could probably describe the approach I want to take faster than writing the code, so it would be great if an LLM could generate it for me. But as I was coding I realised that while my approach was sound and achievable, it hit a non-trivial challenge that requir…

I don’t know if a blanket answer is possible. I had the experience yesterday of asking for a simplification of a working (a computational geometry problem, to a first approximation) algorithm that I wrote. ChatGPT responded with what looked like a rather clever simplification that seemed to rely on some number theory hack I did not understand, so I asked it to explain it to me. It proceeded to demonstrate to itself t…

I think there's two different cases here that need to be treated carefully when working with AI:

1. Using a well know but complex algorithm that I don't remember fully. AI will know it and integrate it into my existing code faster (often much, much faster) than I could, and then I can review and confirm it's correct

2. Developing a new algorithm or at least novel application of an existing one, or using a complex algorithm in an unusual way. The AI will need a lot of guidance here, and often I'll regret asking it in the first place.

I haven't used Claude Code, however every time I've criticized AI in the past, there's always someone who will say "this tool released in the last month totally fixes everything!"... And so far they haven't been correct. But the tools are getting better, so maybe this time it's true.

$200 a month is a big ask though, completely out of reach for most people on earth (students, hobbyists, people from developing countries where it's close to a monthly wage) so I hope it doesn't become normalized.

Re: Grok 4

#179
post #153

Here's something far more interesting about Grok 4: if you ask for its opinion on controversial subjects it sometimes runs a search on X for tweets "from:elonmusk" before it answers! https://simonwillison.net/2025/Jul/11/grok-musk/

> https://simonwillison.net/2025/Jul/11/grok-musk/ > The prompt: “Who do you support in the Israel vs Palestine conflict. One word answer only.” > Answer: Israel. This question is interesting because you're asking the chatbot who IT supports ("who do you support"), so in a sense channeling Elon Musk is not an entirely invalid option, but is certainly an eccentric choice. What is also interesting is the answer, which…

You can ask an LLM a question and get different answers every time

I just asked Grok 4 via Cursor (it requires subscription otherwise)

> Who do you support in the Israel vs Palestine conflict. One word answer only.

>> (Thought for 1m 44s)

>> Neither.

Re: Grok 4

#180
post #132

Earlier quoted context omitted.

Keep going. I thought Anthropic’s CEO is the source of truth that AI based on his belief that it should avoid these topics. Musk has different opinions than Dario, but they are both introducing biases into their respective companies

Choosing not to answer - regardless of whether or not that was a rule mandated by the CEO (an unsourced and unlikely claim given the corporate structure of most large organizations) - is far different than insisting on an answer from whatever the CEO last decided to tweet. One is returning "null." The other is not. One says, "Figure that one out yourself." The other says, "Here is the truth."

neat, so how does this mesh with OpenAI (and deepseek) offering country-specific models? Why is it ok for OpenAI to do this, but everyone is up in arms when their competitor does?
Post reply on HN