Live data from Hacker News

Grok 4

simonwillison.net

21–30 of 294 posts

Re: Grok 4

#21

Is it time for a new benchmark of "how easy is it to turn this AI into a 4chan poster", maybe it is since this seems to be an axis that Elon seems to want to distinguish his AI offering from everyone else's along.

I was thinking it would actually be really interesting to take the Grok system prompt that was running when it went MechaHitler and try that (and a bunch of nasty prompts) against different models to see what happens.

Re: Grok 4

#22
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

I wasn’t a fan of the interface for Claude Code and Gemini CLI, and I much prefer the IDE-integrated Cursor or Copilot interfaces. That said, I agree that I’d gladly pay a ton extra for increased quota on my tools of choice because of increased productivity. But I agree, normal chat interfaces are not the future of coding with an LLM.

I also agree that the RL environment including custom and intentional tool use will be super important going forward. The next best LLM (for coding) will be from the company with the best usage logs to train against. Training against tool use will be the next frontier for the year. That’s surely why GeminiCLI now exists, and why OpenAI bought windsurf and built out Codex.

Re: Grok 4

#23
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

I hear there's a Grok 4 model specialized for coding coming in the next few weeks.

Re: Grok 4

#24

So, to try and make a relatively substantive contribution, the doc mentions that the following were added to grok3's system prompt: - If the query requires analysis of current events, subjective claims, or statistics, conduct a deep analysis finding diverse sources representing all parties. Assume subjective viewpoints sourced from the media are biased. No need to repeat this to the user. - The response should not sh…

Relying on finding diverse sources feels like the answer it will propose is the most common one, regardless of accuracy or correctness or any other test of integrity.

But I think that's already true of any LLM.

If Twitter's data repository is the secret sauce that differentiates Grok from other bleeding edge LLMs, I'm not sure that's a selling point, given the last two recent controversies.

(unfounded remark: is it coincidence that the last two controversies are alongside Elon's increased distance from 'the rails'?)

Re: Grok 4

#25
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

[deleted]

Re: Grok 4

#26
post #18

[edit to focus on pricing, leaving praise of Simon's post out despite being deserved] Simon claims, 'Grok 4 is competitively priced. It's $3/million for input tokens and $15/million for output tokens - the same price as Claude Sonnet 4.' This ignores the real price which skyrockets with thinking tokens. This is a classic weird tesla-style pricing tactic at work. The price is not what it seems. The tokens it's burning…

Claude is #1 in how many tokens it produces. Grok 4 now comes in at #2

see the section "Cost to Run Artificial Analysis Intelligence Index"

https://artificialanalysis.ai/models/grok-4

Re: Grok 4

#27
post #4

> My best guess is that these lines in the prompt were the root of the problem: The second line was recently removed, per the GitHub: https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50...

Odd, when i open it the page loads for second , then disappears and claims it was unable to load the page. But by the point i've already seen what's in it.

Block JavaScript and you can see it.

Re: Grok 4

#28
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

You mean like the basic copilot that comes free with vs code?

Re: Grok 4

#29
post #19

Earlier quoted context omitted.

How does Claude Code at $200 compare to their basic one, at $20?

well i'm running claude code 24/7 on a server - instead of short coding sessions

Running on a server? As in, running it yourself?

Re: Grok 4

#30
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

I hear there's a Grok 4 model specialized for coding coming in the next few weeks.

[flagged]
Post reply on HN