Live data from Hacker News

Grok 4.6

x.ai

511–520 of 696 posts

Re: Grok 4.6

#511
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

* Also FSD is coming this year

Re: Grok 4.6

#512
post #493
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise. So what you experience as a "near simultaneous" release is just the…

This explanation does so much without leaning into conspiracy that the labs are already sitting on the secret sauce but diluting it for the public or being left mystified when a lab drops out of the race for SoTA

Re: Grok 4.6

#513

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

The US govt trusts SpaceXAI for defense and high security missions. The idea they are lying about contracted AI services is absurd. They're also a public company which beings even more oversight than openai / anthropic.

SpaceXAI has huge incentives for not reneging on its commitments to the U.S. government, and those incentives do not exist for entities that lack the power of the purse and guns of the U.S. government.

Furthermore there are plenty of examples of the Trump administration contracting for millions/billions of dollars with companies that aren’t at the top of their game. Are Intel’s fabs best in class because the U.S. bought 10% equity? Are Trump hotels and resorts the best in class because the government expenses for its employees to stay there?

Re: Grok 4.6

#514
post #408
post #406

Does really well and ~2x cheaper than Qwen3.8 2.4T, they have same pricing but grok is around 2x more token efficient: https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/x-ai...

Grok 4.6 vs Sol 5.6 vs Opus 5: https://aibenchy.com/compare/openai-gpt-5-6-sol-low/x-ai-gro...

So Grok 4.6 is incredibly expensive in that comparison? It's double the cost of Sol 5.6 Low and still nearly double of the cost of 5.6 Medium (which scores +6% over Grok).

Re: Grok 4.6

#515

Earlier quoted context omitted.

[flagged]

> So it is still going on How would you know? > Just last week they were fighting Minnesota's law that makes creating this stuff illegal. What law, and what evidence of fighting; and what evidence that their motivation has anything to do with what you allege?

https://news.ycombinator.com/item?id=49105411

Nice semi-colon. Written by AI?

Re: Grok 4.6

#517
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

As great LLMs are, they are no where close to any biological brain. We are not even close to replicating human brain or even brain of an animal. Let’s not add more fuel into this hype.

Re: Grok 4.6

#518
post #493
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise. So what you experience as a "near simultaneous" release is just the…

That’s not how training pipelines work, and would be extremely wasteful for the biggest cost center as well.

Re: Grok 4.6

#519
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

Others have said this too but LLMs are the best approximation of magic we have.

We etch runes on stones, put electricity through them and then try to “convince” them to do our bidding. The answers vary wildly sometimes depending on minutiae.

Prompts should be really called spells. It really feels more like “should I add the frog’s eye or leg into the cauldron” than engineering.

Re: Grok 4.6

#520
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I think a lot of it is just time. The quality of a model is E * C

Where: E = Efficiency, and efficiency gains come from quality of data, quality of algorithms. C = Compute (Size of model, flops of train run)

So a better company can train a bigger and better model with less required compute which let's anthropic get there first. If another company does the same thing with a worse: model architecture, kernel, optimizer, etc... They will get there as well if they just run there train run with more flops for longer

Mythos was actually ready about 6 months ago. So if you have 6 months later or hardware setup and time to train you can get a lot done.

Post reply on HN