Live data from Hacker News

Grok 4

simonwillison.net

81–90 of 294 posts

Re: Grok 4

#85

So, to try and make a relatively substantive contribution, the doc mentions that the following were added to grok3's system prompt: - If the query requires analysis of current events, subjective claims, or statistics, conduct a deep analysis finding diverse sources representing all parties. Assume subjective viewpoints sourced from the media are biased. No need to repeat this to the user. - The response should not sh…

Relying on finding diverse sources feels like the answer it will propose is the most common one, regardless of accuracy or correctness or any other test of integrity. But I think that's already true of any LLM. If Twitter's data repository is the secret sauce that differentiates Grok from other bleeding edge LLMs, I'm not sure that's a selling point, given the last two recent controversies. (unfounded remark: is it c…

Gemini had an aborted launch recently. The controversy there was inserting too much leftist ideology to the point of spewing complete bs.

Re: Grok 4

#86
post #80
post #73

Earlier quoted context omitted.

I don’t know if a blanket answer is possible. I had the experience yesterday of asking for a simplification of a working (a computational geometry problem, to a first approximation) algorithm that I wrote. ChatGPT responded with what looked like a rather clever simplification that seemed to rely on some number theory hack I did not understand, so I asked it to explain it to me. It proceeded to demonstrate to itself t…

You really can’t compare free "check my algorithm" ChatGPT with $200/month "generate a working product" Claude Code. I’m not saying Claude Code is perfect or is the panacea but those are really different products with orders of magnitude of difference in capabilities.

Claude 4? Or is Claude Code really so much better than say Aider also using Claude 4?

Re: Grok 4

#87
post #4

> My best guess is that these lines in the prompt were the root of the problem: The second line was recently removed, per the GitHub: https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50...

How do you even QA the non-determinism of these technologies?

Re: Grok 4

#90
post #43

I didn't follow the Mechahitler issue can someone explain the technical reasons that it happened? Was grok4 released early or was there a variant model used for @grok posts that's separate from grok4?

It was grok 3, and it was tricked/prompted to reply like so, just like any other LLM can be. Apparently at one point it was prompted with a choice between identifying itself as a MechaHitler or a GigaJew, so it chose the former.
Post reply on HN