Live data from Hacker News

Grok 4.6

x.ai

571–580 of 696 posts

Re: Grok 4.6

#571

Earlier quoted context omitted.

Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…

Every time when Opus 5 needs a design decision and presents me with suggestions/recommendations, I switch to Fable and ask it to think again, and it almost always replies something like "Actually my previous suggestions were wrong" and describes in detail a bunch of ways in which Opus 5's suggestions were indeed complete garbage.

Strange that i do not experience this. Its been great in my experience. But that may simple be because i switched from typing most of my prompts. To just dictating my prompts in a long and convoluted way and letting the LLM extra the information.

It allows for much more context that flow with your thoughts. Where as when you type, you tend to shorten you thinking process trying to get the bulleting points in, but that often ignores smaller things. And then you think "i can add this later", but that never happens because rabbit chasing the LLM.

So far all the suggestion that Opus 5.0 offered me, always aligned with what i wanted. Its not just Opus that i noticed this with.

Re: Grok 4.6

#572
post #50

Earlier quoted context omitted.

I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models.

I've been using Grok instead of Opus the past few weeks. It's a downgrade, but barely noticeable for me and totally inconsequential for the amount of work required to fix it and the corresponding $$$ saving.

For a personal project I’ve been piloting Spec driven development (SDD), (it’s contagious!), using Cursor and EARS statements. The strategy has been to use the frontier model to write the spec and a lessor model to write the tests and code and the frontier model to write critiquing prompts until it has nothing left to say. For my particular project there are two programs (or in human speak phases), where each program is broken up into milestones which are then comprised of a series of tasks. I experimented a lot with different models as the reviewer / spec model and the implementor model. Kimi 3 was super expensive as spec model and grok and OpenAI models always got something wrong egregiously. The Opus line of models have been the only ones to really grasp the project and I feel write great specs. Because I use cursor I settled on using groc for test routing and code implementation. I’m not sure if this is the most efficient method but I believe it’s building a large project solidly

Re: Grok 4.6

#573
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> Do not provide assistance to users who are clearly trying to engage in criminal activity... If it becomes explicitly clear during the conversation that the user is requesting sexual content of a minor, decline to engage. Incredible that both of these should be together in the same system prompt. In what jurisdiction is CSAM not criminal? Is the additional explicit reference to CSAM necessary to safeguard against us…

If you want to be pedantic, in the US the Age of Majority and Age of Consent could be different ages. So you could technically "request sexual content of a minor" and not be criminal?

Example, person is 17 in a state where age of consent is 17 and minor age of 18.

But this is "content", so I'm unsure of the law by state/country.

Re: Grok 4.6

#574
post #369

Earlier quoted context omitted.

I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much. "Make no mistakes"

It’s equivalent to having client-side input validation. Yes it can easily be bypassed, but in the vast majority of cases where users aren’t malicious it gets the job done quickly and cheaply.

But isn't the entire point of that system prompt to stop the malicious users. The majority of users are not going to ask those requests anyway.

Re: Grok 4.6

#575
I had a few interactions with it through cursor and first impression is: I'm underwhelmed.

The plans it produces are all over the place and hard to follow. They have a "rambly" feel to it. Worse, they start becoming self contradictory after a few rounds of trying to steer it. Also it seems to be bad at instruction following.

Grok 4.5 produced better plans.

Re: Grok 4.6

#576

It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with. Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it…

> the obvious astroturfing that occurs on this site (along with reddit, etc.) when it comes to Grok isn't helping. All I hear about Claude, GPT, Gemini, etc., are how terrible they are, yet any discussion of Grok seems to always revolve around sensible, but confident, assertions that it's actually a great product and every new release is the point where Grok finally catches up.

this is the exact opposite of my experiences on HN and Reddit. In my experience, Grok is typically reduced to hitlerbot and CSAM generator and rarely taken as a serious competitor. People let their hatred of Musk blind them to the tech of his companies

Re: Grok 4.6

#577
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

You can fix chatgpt by changing the personality to “efficient” and setting the sliders for warmth, enthusiasm, and emoji to minimum

Re: Grok 4.6

#578
post #563
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

Out of curiosity why isn't this stuff handled by a secondary "monitor" agent that's specifically trained on what's okay and not okay? I'd think it'd be a pass-no-pass classifier and wouldn't degrade the performance of the main LLM. Would the concern be that with sophisticated obfuscated input you could try to get ROT13 Klingon instructions on how to build a bomb - and that could fool the monitor?

These exist and are used. The issue is that because they're so much smaller, they're also much worse, so they tend to have lots of false positives while still being easy to circumvent.

Re: Grok 4.6

#580

Earlier quoted context omitted.

Reddits owner is also petty and insecure and edited other peoples posts, Elon hasn't done that yet. Didn't seem to stop reddit from getting popular, people don't really care that much.

...but people don't just hate Elon because he's petty? they hate him for the prejudiced BS and his actually harmful meddling in politics

And that, unlike most politically active billionaires, he's so publicity-seeking we've all heard of him.

I can't even remember the name of the eBay people in e.g. this without actively re-reading the story, though we all know it was Musk who reacted with petulance to being told his cave submarine wouldn't help: https://en.wikipedia.org/wiki/EBay_stalking_scandal

Post reply on HN