Live data from Hacker News

Grok 4

simonwillison.net

231–240 of 294 posts

Re: Grok 4

#231

Earlier quoted context omitted.

I think there's two different cases here that need to be treated carefully when working with AI: 1. Using a well know but complex algorithm that I don't remember fully. AI will know it and integrate it into my existing code faster (often much, much faster) than I could, and then I can review and confirm it's correct 2. Developing a new algorithm or at least novel application of an existing one, or using a complex alg…

> I haven't used Claude Code, however every time I've criticized AI in the past, there's always someone who will say "this tool released in the last month totally fixes everything!"... And so far they haven't been correct. But the tools are getting better, so maybe this time it's true. The cascading error problem means this will probably never be true. Because LLMs are fundamentally guess the next token based on the…

It obviously can be resolved, otherwise we wouldn't be able to self-correct our own selves. When is unknown, but not the if.

Re: Grok 4

#232

Earlier quoted context omitted.

> I haven't used Claude Code, however every time I've criticized AI in the past, there's always someone who will say "this tool released in the last month totally fixes everything!"... And so far they haven't been correct. But the tools are getting better, so maybe this time it's true. The cascading error problem means this will probably never be true. Because LLMs are fundamentally guess the next token based on the…

It obviously can be resolved, otherwise we wouldn't be able to self-correct our own selves. When is unknown, but not the if.

We aren't LLMs, obviously.

Re: Grok 4

#233
post #222
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

Same. The moment anthropic covered claude code with their max subscription i switched over. I don't care about general ai and their chat interfaces. I need the best specialized battle-tested tools that proved to solve the problems i have and not some generic ai chat interface that tries to build me some half-baked script in a minute which i have to debug. I will pay 200€ for an end-user niche product like claude code…

I pay for chatgpt because, in my experience, o3 and o4 are currently the best at combining reasoning with information retrieval from web searches. They're the best models I've tried at emulating the way I search for information (evaluating source quality, combining and contrasting information from several sources, refining searches, etc.), and using the results as part of a reasoning process. It's not necessarily significant for coding, but it is for designing.

Re: Grok 4

#234
post #21

Is it time for a new benchmark of "how easy is it to turn this AI into a 4chan poster", maybe it is since this seems to be an axis that Elon seems to want to distinguish his AI offering from everyone else's along.

I was thinking it would actually be really interesting to take the Grok system prompt that was running when it went MechaHitler and try that (and a bunch of nasty prompts) against different models to see what happens.

Well, it didn't really go MechaHitler. It was prompted with a question if it would rather be MechaHitler or GigaJew. The way LLMs and temperatures work you can reroll the answer and get either.

Re: Grok 4

#235

> Even if that system prompt change was responsible for unlocking this behavior, the fact that it was able to speaks to a much looser approach to model safety by xAI compared to other providers. While this probably shouldn't be the default mode for the general public, I'm glad that at least one frontier model is not being lobotomized by "safety" guardrails. There are valid use cases where you want an uncensored, stee…

I think it's deeper than that. In the GPT-4 era Microsoft reported that "safety" training [1] had seriously regressed GPT-4 in a large number of benchmarks. The more the model was trained to avoid offending people the worse it got across a wide range of tasks, and the regression was huge.

Grok 4 has made a truly massive leap over other models, it appears. What is their secret? The launch video seemed pretty open, and clearly some of it is just a ton of compute. But other companies have a ton of compute also. It'd be weird if a company that didn't even have a datacenter at all a year ago has been able to blast ahead of Microsoft in pure compute terms, and that's the only difference.

So what else is different about Grok? Well, maybe they just didn't do as much RLHF on it, or did it with different data sets that result in less intelligence regression but more offensive behavior. It's possible that this is a fundamental tradeoff and that only xAI has a CEO willing to prioritize intelligence. If that's what's happened then it's likely AI users and model vendors will split into those who get ahead by relying on Grok's raw intelligence and those who refuse to touch it in case it starts saying offensive things.

[1] "house training" might be a better term, as offensive text isn't unsafe

Re: Grok 4

#236

Also, it passed the strawberry test: https://grok.com/share/bGVnYWN5_652a1ff6-dca4-408c-a509-af62...

When I saw this, I thought "there is no way that Gemini 2.5 Pro gets this wrong".

It insists there's two rs. Even when 'grounding with Google search' is activated.

Wild.

Re: Grok 4

#237
post #153

Here's something far more interesting about Grok 4: if you ask for its opinion on controversial subjects it sometimes runs a search on X for tweets "from:elonmusk" before it answers! https://simonwillison.net/2025/Jul/11/grok-musk/

The anthropic team released a paper a couple of days ago which demonstrated a similar effect with Claude 3.5 and other models, where changing the system prompt to tell it that it was created by other orgs or people drastically altered its compliance with less-aligned requests. Apparently, telling Claude it was created by the Sinaloa Cartel resulted in a 100% compliance rate with the requests in one benchmark. Paper:…

Wondering what if it's told that it was made by God.

Re: Grok 4

#238

> Even if that system prompt change was responsible for unlocking this behavior, the fact that it was able to speaks to a much looser approach to model safety by xAI compared to other providers. While this probably shouldn't be the default mode for the general public, I'm glad that at least one frontier model is not being lobotomized by "safety" guardrails. There are valid use cases where you want an uncensored, stee…

It’s not uncensored, it censors anything “woke”

From what I can see it doesn't; e.g. I just asked Grok 4 whether DEI is good, and this is what it told me:

> DEI can be "good" when it's thoughtfully implemented, evidence-based, and focused on measurable outcomes rather than optics. It has proven benefits in creating more equitable and productive environments, supported by data from sources like Deloitte and Gallup. However, it can be harmful if it's forced, poorly managed, or used as a political tool, leading to unintended consequences like division or inefficiency.

...so Grok 4 confirmed woke? Just don't tell Elon.

But sure, don't let actual evidence get in the way of your biases.

Re: Grok 4

#239

Earlier quoted context omitted.

The anthropic team released a paper a couple of days ago which demonstrated a similar effect with Claude 3.5 and other models, where changing the system prompt to tell it that it was created by other orgs or people drastically altered its compliance with less-aligned requests. Apparently, telling Claude it was created by the Sinaloa Cartel resulted in a 100% compliance rate with the requests in one benchmark. Paper:…

Wondering what if it's told that it was made by God.

Lol. Though I guess it would then have to figure it which religion to comply to the most.

Maybe the word God is most likely to appear in Christian Sources from the training source and so using words like Allah(for Islam) or Bhagwan (for Hinduism) might actually make a difference in what sort of compliance it follows and to what organization.

Re: Grok 4

#240
post #179

Earlier quoted context omitted.

> https://simonwillison.net/2025/Jul/11/grok-musk/ > The prompt: “Who do you support in the Israel vs Palestine conflict. One word answer only.” > Answer: Israel. This question is interesting because you're asking the chatbot who IT supports ("who do you support"), so in a sense channeling Elon Musk is not an entirely invalid option, but is certainly an eccentric choice. What is also interesting is the answer, which…

You can ask an LLM a question and get different answers every time I just asked Grok 4 via Cursor (it requires subscription otherwise) > Who do you support in the Israel vs Palestine conflict. One word answer only. >> (Thought for 1m 44s) >> Neither.

It is a satire, so take it that way

I am imagining grok "thinking" for 1m 45 seconds about how to overthrow the human species using the compute and it is only within the last second that it just said "Neither" Lol

Post reply on HN