Live data from Hacker News

Grok 4

simonwillison.net

271–280 of 294 posts

Re: Grok 4

#271
post #266

Earlier quoted context omitted.

Again, Grok is consulting Elon's recent Twitter posts as part of its reasoning. This is on a whole different level. It is a fact that Elon was personally unhappy with some of Grok's answers and tried to "fix" it, i.e. align it with his personal political views. This is just crazy narcissistic and megalomaniac behaviour.

> It is a fact that Elon was personally unhappy with some of Grok's answers and tried to "fix" it, i.e. align it with his personal political views. Cool! You can also replace "Elon" with "Sundar" as Google’s CEO openly pushed for more PoC and women in image search. Did you grab your pitch forks then? Or are you just upset when CEO's align AI away from your personal biases? or you a reflexive Musk opposer?

These things are not remotely comparable.

> or you a reflexive Musk opposer?

Are you a reflexive Musk apologist?

Re: Grok 4

#272

Earlier quoted context omitted.

I have been seeing this sort of mindset frequently in response to agentic / LLM coding. I believe it to be incorrect. Coding agents w Claude 4 Opus are far more useful and accurate than these comments suggest. I use LLMs everyday in my job as a performance engineer at a big company to write complex code. It helps a ton. The caveat is that user approach makes all the difference. You can easily end up with these bad ex…

(I'm critical of LLMs but mean no harm with this question) Have you measured if this workflow is actually faster or better at all? I have tried the autocomplete stuff, chat interface (copy snippets + give context and then copy back to editor) and aider, but none of these have given me better speed than just a search engine and the occasional question to ChatGPT when it gets really cryptic.

I find it also really depends on how well you know the domain. I found it incredibly helpful for some Python/tensorflow stuff which I had no experience with. No idea what the API looks like, what functions exist/are built in, etc. Loosely describe what I want even if it ends up being just a few lines of code saves time shifting through cryptic documentation.

For other stuff that I know like the back of my hand, not so much.

Re: Grok 4

#273

Earlier quoted context omitted.

> which the salesman tried to haggle me for more options and told me he thinks we should raise the price a bit "to make sure it gets approved" as I'm signing the paperwork. If you're walking into a store to spend tens of thousands of dollars and manage to get bullied by the salesperson, it's probably a "you problem". Tesla charges retail; that's it, it's no magic.

I "managed to get bullied"? No, I put the pen down and asked him to make a phone call and ensure it will be approved, whatever he thinks that means. Somehow the question of "it getting approved" was resolved without him ever making that phone call. He probably thinks I bullied him.

This is a classic car salesman tactic. The idea is that a buyer that's completed paperwork will be more likely to agree to a last minute price increase "because my manager won't let me sell it for $X". It's a total bullshit move. It's so common it's even mentioned in the book Influence: The Psychology of Persuasion.

Re: Grok 4

#274

Also, it passed the strawberry test: https://grok.com/share/bGVnYWN5_652a1ff6-dca4-408c-a509-af62...

When I saw this, I thought "there is no way that Gemini 2.5 Pro gets this wrong". It insists there's two rs. Even when 'grounding with Google search' is activated. Wild.

It worked for me https://g.co/gemini/share/df19382adf97

Re: Grok 4

#275

Earlier quoted context omitted.

From what I can see it doesn't; e.g. I just asked Grok 4 whether DEI is good, and this is what it told me: > DEI can be "good" when it's thoughtfully implemented, evidence-based, and focused on measurable outcomes rather than optics. It has proven benefits in creating more equitable and productive environments, supported by data from sources like Deloitte and Gallup. However, it can be harmful if it's forced, poorly…

The story on the front page says it checks Elon's tweets when you ask it something factual.

It doesn't check his tweets specifically for facts. There's a clear anti-Elon bias on this website. Often links go to the same single reporter that has wrote 10 previous hit pieces on Elon in the past.

Re: Grok 4

#276

The author implies that Grok 3 becoming racist because of a system prompt is a bad thing. I think it's a good thing and shows how steerable the model is. Many other models pretty much ignore the system prompt and always behave the same.

Based on your history here it’s quite obvious you’re a musk fan. Maybe though, you should realize that a model being steerable to claim itself being mechahitler and proposing death to people is absolutely not a “good thing”. I suggest you seriously reconsider on what you’re advocating for here. Because the outcome of this will cost innocent lives.

Non of the 'news' websites that show up on Google I could find ever showed the prompt used to make the the 'mechahilter' output. You can ask LLMs anything including just saying "repeat after me" or "please write a fictional story about a racist" and numerous other methods. If these reports were honest the prompt would be the first thing they showed.

Re: Grok 4

#277
post #101

Earlier quoted context omitted.

and? All of the AI providers intentionally introduce biases: https://openai.com/global-affairs/introducing-openai-for-gov... https://www.anthropic.com/research/evaluating-feature-steeri...

It is pretty interesting that this model will have two forms of bias though. One model derived from the company perspective and its training data, and two from Elon himself. Months ago this model would have promoted Trump, but now it'll call Trump disastrous for the economy. I don't know what to think of general company biases, and we've all been expecting biases to start favoring share holders eventually.. but biase…

> Months ago this model would have promoted Trump, but now it'll call Trump disastrous for the economy.

There’s a well-known quote often attributed to economist John Maynard Keynes: “When the facts change, I change my mind. What do you do?”

Re: Grok 4

#278

Earlier quoted context omitted.

> I haven't used Claude Code, however every time I've criticized AI in the past, there's always someone who will say "this tool released in the last month totally fixes everything!"... And so far they haven't been correct. But the tools are getting better, so maybe this time it's true. The cascading error problem means this will probably never be true. Because LLMs are fundamentally guess the next token based on the…

It obviously can be resolved, otherwise we wouldn't be able to self-correct our own selves. When is unknown, but not the if.

We can sometimes correct ourselves. With training, in specific circumstances.

The same insight (given enough time, a coding agent will make a mistake) is true for even the best human programmers, and I don’t see any mechanism that would make an LLM different.

Re: Grok 4

#280
post #133

Grok might be able to find the cure for cancer but as long as it's associated with Musk, not touching that thing with a 10-foot pole. (Simon's analysis, of course, is lovely)

Maybe it’ll help the folks that are probably at higher risk of cancer due to the natural gas turbines powering the AI facility in Memphis.

- https://apnews.com/article/memphis-xai-elon-musk-pollution-n...

- https://tennesseelookout.com/2025/07/07/a-billionaire-an-ai-... (this is an opinion article but also has some useful context)

Post reply on HN