Live data from Hacker News

Grok 4.6

x.ai

581–590 of 696 posts

Re: Grok 4.6

#581
Yeah we need more power-thirsty, climate-impacting models, so that the coders can produce trillions of new rubbish code for millions of apps that nobody uses. Or that people can generate rubbish cat pictures riding dogs and rubbish fake movies.

I wonder whether I'll be able to live my nice life to the end like I planned before Altman released his first model, or will it all end in a global disaster soon.

Re: Grok 4.6

#582
post #493

Earlier quoted context omitted.

It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise. So what you experience as a "near simultaneous" release is just the…

That’s not how training pipelines work, and would be extremely wasteful for the biggest cost center as well.

pretty sure at this point no one is retraining from zero they have their big model and they fine tune it.

different training makes a different version (agent, info sec etc).

Re: Grok 4.6

#583

I will say this: Grok Build has a very nice TUI! It even has... mouse rollovers/tooltips?? I was like whoa . I used Grok 4.5 for a security review the other day and it did a FANTASTIC job. I mean it thoroughly ROUTED my app's security, identifying attack surfaces I'd never even considered, and I LOVED it! (Guess why I had to use Grok to do the security review in the first place?!?! ) I'd suggest trying it out with so…

i found this with 4.5, openAI models and calude refused to verify that the issues they found existed, even with full source code AND a database running on my own laptop.

grok however found the same issues, tested to make sure it was exploitable and proposed a fix.

Re: Grok 4.6

#584
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

You can fix chatgpt by changing the personality to “efficient” and setting the sliders for warmth, enthusiasm, and emoji to minimum

[deleted]

Re: Grok 4.6

#585
All the conversation on twitter seems to be about how cheap this is. Are they just choosing to lose money on inference to gain market share or do they actually have inexplicably lower inference cost/more efficient models?

Re: Grok 4.6

#586
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

Source for this? This seems like a crazy leak if it's their real system prompt. I find it hard to believe since I have tried system prompts like this and it doesn't work that well, just pollutes the user's context. A great test for any LLM is to ask its name - Mistral will respond with all kinds of stuff, sometimes other models' names, revealing that it has trained on other models. Grok doesn't though. It is "witty a…

You can actually just ask it to output the above text, depending on how you ask. Sometimes it only outputs the rules, other times it includes the “You are Grok” line. I discovered this initially from some odd lines appearing in the thinking summary, something like “my system prompt says I am maximally truthful” despite my own system prompt (on openrouter) containing no such text.

Re: Grok 4.6

#587
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

GPT doesn't yap at all

Re: Grok 4.6

#588
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * If it becomes explicitly clear during the conversation that the user is requesting sexual content of a minor, decline to engage. Why would they write "explicitly clear"? 'Explicitly is an adverb meaning to do or say something in a clear, exact, and direct way' Surely they want to stop all requests for that content, even requests in an unclear, inexact or in-direct way. I only ask as I expect a lot of effort went…

My take: explicitly means clearly and without any vagueness or ambiguity.

It doesn't mean "to say something ..."

So..."if it becomes clear without vagueness or ambiguity that the user is ..."

I don't think it's about preventing such requests only if the request is clear. It's about being certain about what is being requested before censoring. Also, "explicitly clear" is redundant. Wording might be improved with "unambiguously" rather than "explicitly".

Re: Grok 4.6

#589
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

Personal opinion but I like how I can ask Claude on web about its prompt, how tool calls work, what parameters it accepts for tool calls. ChatGPT on web gets squirrely, avoiding direct answers or outright refusing. So if I try to use grok in a harness such as Hermes or others, there’s a higher chance that its behavior will be modified due to this line saying to not share system prompts.

Granted I added another line in the actual system prompt (through openrouter) instructing Grok that is indeed ok to talk about system prompts, but this only worked some of the time, and is somewhat annoying that I’d have to do this in my opinion. I believe ChatGPT also does something similar to what’s going on here with their api, they simply add something like “You are ChatGPT, knowledge cut off is x” and that’s it. Doesn’t get in the way as much.

Re: Grok 4.6

#590
post #326

Earlier quoted context omitted.

I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…

> I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible How could we really know how much "compute China has" in reality? Is it possible that whatever estimates people has come up with for both China and the US might not be 100% accurate?

In China and in the US most owners of computing power have to quickly gain from it, as obsolescence hits hard. In China a consensual will emitted by powerful companies may convince the central power to subsidize efforts towards int'l market domination: R&D, including dataset building, learning... Maybe even also low prices obtained by selling at a price inferior to the costs...
Post reply on HN