Live data from Hacker News

Grok 4.6

x.ai

641–650 of 696 posts

Re: Grok 4.6

#641
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

GPT doesn't yap at all

I don't think it's as significant anymore but there was a point where GPT would write 8 paragraphs to say yes, while Grok would literally just say "Yes." They've converged a bit (Grok now says more, GPT says less) but I'd argue Grok is still more efficient while Claude/GPT are more wordy (and often needlessly)

Re: Grok 4.6

#642

Earlier quoted context omitted.

> So it is still going on How would you know? > Just last week they were fighting Minnesota's law that makes creating this stuff illegal. What law, and what evidence of fighting; and what evidence that their motivation has anything to do with what you allege?

https://news.ycombinator.com/item?id=49105411 Nice semi-colon. Written by AI?

> Nice semi-colon.

Thanks. I've been using them literally for decades. This one is a deliberate stylistic alteration, where a comma would have been the obvious choice, intended to suggest a specific speaking cadence.

Your source directly refutes your argument:

> In the 38-page lawsuit, xAI — whose AI model chatbot and image generator Grok is available on the social media platform known as X, formerly Twitter, and elsewhere — said it does not contest the state’s interest in banning the distribution of AI-generated nude images of real people without their consent. But it said Minnesota’s law “extends far beyond that goal,” banning many constitutionally protected images and video and subjecting the company to a penalty of $500,000 per violation.

Re: Grok 4.6

#643
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

To add to that: given what they went through with the last model, I don't believe for a second that the real system prompt is even remotely this short.

Re: Grok 4.6

#644

Earlier quoted context omitted.

The alternative is Claude-style "safeguards" aka censorship, which: 1. doesn't eliminate the possibility of a jailbreak anyway 2. frequently has false positives, triggering on innocuous requests, which is just really annoying Not saying that we can't (or shouldn't) do better than Grok, but I really don't know what the best solution is here...

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

Can't you do that with any model you can run on your own hardware ?

If you rent other people's shit can't be surprised when they have restrictions on what you can do with it. I would guess renting a car comes with some similar clauses

Re: Grok 4.6

#645
post #595

Earlier quoted context omitted.

Is there concrete evidece that those are xAI's default prompts anyway? They seem plausible enough but how would company outsiders know?

I've spent the last year working as an annotator/evaluator for DataAnnotation. All the frontier/flagship model providers use independent contractors for iterating on their LLMs. I'm not able to tell you which models I've worked on as a term of my NDA. The system prompt seems plausible, but in my experience they are much much much much longer and more verbose.

What layer do those typically work at and how ? And how did you get into the field ?

Re: Grok 4.6

#646
post #379

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

Given how much we know of the human brain this may be partially how we work

Maybe we should be asking what our own "system prompts" are ?

Re: Grok 4.6

#647
post #58

Earlier quoted context omitted.

In my tests Grok 4.5 is definitely not Opus level. It is somewhere in between Sonnet and Opus, I'd say maybe a bit closer to Sonnet. We'll see with 4.6.

In my experience Grok 4.5 codes at Opus 4.8 level, and being much faster as cheaper, I can just ask it to do self-review and the final reviewed code is _better_ than Opus 4.8 for the same time/budget. But Opus 5/4.8 was better for non-code architecture discussions and general intelligence. However, for the cost, I'd use GPT 5.6 Sol and get much better results. Interestingly, Sol is not great for coding - slow and ove…

You're surprised that the model reviewing itself thinks its code is better?

Re: Grok 4.6

#648
Kinda good, but burns a ton of tokens compared to sol. Gave it mid sized task, it burned entire context on thinking and then compacted after first edit. Sol would use 20-30% of context in comparison

Re: Grok 4.6

#649
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

There is a herd of companies all running a race. The technology is known. They all have roughly the same resources. It’s not unexpected that they have similar cycle times for model development and that those models will be of roughly the same quality. Then layer in corporate PR demands and you see all these models landing within weeks, sometimes days, of each other to keep the model developer’s name associated with “…

It still feels a bit strange to me that no one comes up with "secret sauce" that gives a big edge for at least a few months.

Re: Grok 4.6

#650

Earlier quoted context omitted.

They do "reflect top to bottom". Hold a written word in front of your eyes to read it. Now flip it to the mirror to read the reflection: Did you flip it horizontally? Then it reflected left-to-right. Did you flip it vertically? Then it reflected top-to-bottom.

I think "reflect top to bottom" is intended to mean "swap top and button". A mirror reflects left, right, top and bottom perfectly. It's front and back that it swaps. Someone saying that a mirror swaps left and right is comparing it to a photograph, and only because we, as bipedal creatures, really prefer to orient images of other humans with heads up.

Someone saying that a mirror is swapped left and right is because they rotated themselves 180 degrees about the vertical axis to face the vertically aligned mirror. If they used a horizontal axis instead, they would have swapped top and bottom. And if the mirror is horizontally mounted on the floor, anything goes. You'd probably say it swaps up and down, which is front and back from the mirror's point of view.
Post reply on HN