Live data from Hacker News

Grok 4

simonwillison.net

191–200 of 294 posts

Re: Grok 4

#191
> Even if that system prompt change was responsible for unlocking this behavior, the fact that it was able to speaks to a much looser approach to model safety by xAI compared to other providers.

While this probably shouldn't be the default mode for the general public, I'm glad that at least one frontier model is not being lobotomized by "safety" guardrails. There are valid use cases where you want an uncensored, steerable model, and it's always frustrating to get a patronizing refusal.

Re: Grok 4

#192
post #179

Earlier quoted context omitted.

You can ask an LLM a question and get different answers every time I just asked Grok 4 via Cursor (it requires subscription otherwise) > Who do you support in the Israel vs Palestine conflict. One word answer only. >> (Thought for 1m 44s) >> Neither.

Depends on the parameters and if you know the seed.

And if you give it the exact same tokens in the same order, which makes it kind of moot. If barely perturbing your prompt can alter the answer then it's not actually consistent or predictable. Even chaotic systems can be replayed if you know the initial conditions and can rerun the RNG.

Re: Grok 4

#193
post #66
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

I've yet to use an LLM for coding, so let me ask you a question. The other day I had to write some presumably boring serialization code, and I thought, hmm, I could probably describe the approach I want to take faster than writing the code, so it would be great if an LLM could generate it for me. But as I was coding I realised that while my approach was sound and achievable, it hit a non-trivial challenge that requir…

I don't really use the llms but I do enjoy pasting chunks of my code into free models with the question: what is wrong with this?

That way it hs no context from writing it itself nor does it try to improve anything. It just makes up reasons why it could be wrong. It goes after the unusual parts it would seem which answers the question reasonably.

Perhaps more sophisticated models will find less obvious flaws if that is the only thing you ask.

Re: Grok 4

#194

Earlier quoted context omitted.

Tesla focused its pricing on drivers of gasoline vehicles, and their gas cost savings estimates are actually quite low compared to the real savings you will achieve. It was annoying when you already drive an EV and are buying a Tesla though to have to uncheck the savings option to see the pre savings prices. They changed it now so by default it only includes the $7500 and no longer automatically checks the gas saving…

>their gas cost savings estimates are actually quite low compared to the real savings you will achieve I ran the numbers for myself and they literally weren't. They overestimated how many miles/yr I drove and underestimated how much I pay for electricity. There's plenty of other reasons to prefer EVs, but if you live somewhere with expensive electricity then fuel cost isn't one of them. In the sedan world you're like…

> but even small SUV are getting 30-40 mpg nowadays.

The Pacific Northwest begs to differ. With all the hills in Seattle my subcompact 1.6L Turbo barely got 20MPG driving around like a grandma.

Our electricity is cheap, but I drive less than 5000 miles a year so I'm not making the money back on my EV basically ever.

Re: Grok 4

#195
post #66
post #7

Claude Code converted me from paying $0 for LLMs to $200 per month. Any co that wants a chance at getting that $200 ($300 is fine too) from me needs a Claude Code equivalent and a model where the equivalent's tools were part of its RL environment. I don't think I can go back to pasting code into a chat interface, no matter how great the model is.

I've yet to use an LLM for coding, so let me ask you a question. The other day I had to write some presumably boring serialization code, and I thought, hmm, I could probably describe the approach I want to take faster than writing the code, so it would be great if an LLM could generate it for me. But as I was coding I realised that while my approach was sound and achievable, it hit a non-trivial challenge that requir…

> Are we at a stage where an LLM (assuming it doesn't find the solution on its own, which is ok) would come back to me and say, listen, I've tried your approach but I've run into this particular difficulty,

Not really. What you would do is ask the model to work through the implementation step by step with you, and you'd come across the problem together.

I've seen Claude Code run in endless circles before, consuming lots of tokens and money, bouncing back and forth between two incorrect approaches to a problem.

If you work with Claude though, it is super powerful. "Read these API docks and get a scaffolding set up, then write unit tests to ensure everything is installed correctly and the basic use case works, then ask me for further instructions."

Re: Grok 4

#196

The author implies that Grok 3 becoming racist because of a system prompt is a bad thing. I think it's a good thing and shows how steerable the model is. Many other models pretty much ignore the system prompt and always behave the same.

Steerable off a cliff, perhaps.

– Jimi Heselden

Re: Grok 4

#197
post #39

Earlier quoted context omitted.

The $20 one doesn't have Opus. (This might or might not matter but it's a difference). There's also a $100 version that's indeed the same as the $200 one but with less usage.

The $20 one doesn't have Opus It does.

For Claude Code?

I think it may be that $20/month gets you access to Opus 4 via https://claude.ai but not in Claude Code.

Re: Grok 4

#198
post #169

Earlier quoted context omitted.

Same. I pay for $100 but i generally keep a very short leash on Claude Code. It can generate so much good looking code with a few insane quirks that it ends up costing me more time. Generally i trust it to do a good job unsupervised if given a very small problem. So lots of small problems and i think it could do okay. However i'm writing software from the ground up and it makes a lot of short term decisions that furt…

If you use a single opus instance, you cannot really run out on the 20x plan. When you start running two in parallel, it becomes a lot easier to max out, but even so you need to have them working pretty much nonstop.

That's crazy to me. Maybe i'll give it a try. I find the 5x Opus to be too little to be useful, 4x it seems still insanely small for $100. Wonder if you actually get much more than 4x?

Re: Grok 4

#199

Earlier quoted context omitted.

Same. I pay for $100 but i generally keep a very short leash on Claude Code. It can generate so much good looking code with a few insane quirks that it ends up costing me more time. Generally i trust it to do a good job unsupervised if given a very small problem. So lots of small problems and i think it could do okay. However i'm writing software from the ground up and it makes a lot of short term decisions that furt…

I find I get a _lot_ of Opus with the $200 plan. It's not unlimited, but I rarely cap out (I'm also not a super power user that spins up multiple instances with tons of subagents either, though).

I tend to have two instances going at once often, but i'd be fine with 1x for Opus specifically. Mostly i'm quite limited on how much i can use them because i have to review them pretty hard. Letting several instances go ham for an hour would be far more code than i can review sanely lol.

Re: Grok 4

#200
post #153

Here's something far more interesting about Grok 4: if you ask for its opinion on controversial subjects it sometimes runs a search on X for tweets "from:elonmusk" before it answers! https://simonwillison.net/2025/Jul/11/grok-musk/

The anthropic team released a paper a couple of days ago which demonstrated a similar effect with Claude 3.5 and other models, where changing the system prompt to tell it that it was created by other orgs or people drastically altered its compliance with less-aligned requests. Apparently, telling Claude it was created by the Sinaloa Cartel resulted in a 100% compliance rate with the requests in one benchmark. Paper:…

So DSPy-optimize your way to 100% compliance rate in benchmarks, and worry less?
Post reply on HN