Live data from Hacker News

Grok 4.6

x.ai

491–500 of 696 posts

Re: Grok 4.6

#491
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

I think it makes sense. You wouldn’t want to hire an employee who’s intellectually incapable of helping customers commit a crime. You’d want to give them instructions, and have them follow their instructions.

Re: Grok 4.6

#492

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

So basically, nothing that actually affects working with it in August 2026. Got it. Facebook has a far longer (and worse) laundry list of offenses and I'm sure you still use it. Or Threads, or Instagram. > My organization has outright banned Grok That's too bad, as it's currently the only model that won't consistently flag honest good-actor security questions, in my experience. So I'd ask you who you work for, but I…

Your experience is not reflective of mine at all, or my colleagues’, so I would check out the better SOTA models out again. I use codex extensively for security-related work - much of which is _overtly_ offensive - without issue. Same for Claude, minus Fable, after going through their approval process. I also went through OpenAI’s, but theirs was just basic KYC and instant. GPT-5.6 in Codex has produced full chain RCEs, ASLR bypass and all, in ubiquitous software with nothing more than a prompt and a few days of crunching. Things you’d be paying $$$ for just last year, now produced on not much more than a whim & a prompt. I can’t speak for Grok’s abilities wrt these types of things, but for your own sake, take the 5 mins it takes to complete the verification processes for OAI/Anthropic if you work in security.

Also, assuming people use Meta/FB/Instagram here, of all places, is certainly an assumption - very poor fodder for a “gotcha”. I find Elon’s political activities and the social beliefs he uses his purchased platform to spread loathsome and daft, and it will take a lot more than “almost as good on benchmarks but cheaper” to let my fiscal tendencies outweigh my moral ones. I’ve held similar beliefs for Zuck for far longer and have cut everything marred by the slime of his tentacles out of my digital life for years, as _many_ here have also done. Accusing someone of uneven application of moral influence over their decisions when you only have information relating to a single decision is poor argumentation.

If you find what Musk spreads palatable, or maintain distance and a lack of awareness, or just don’t care - fine. But don’t confuse the hill you chose with a moral high ground. Any snark you launch from such a position is likely going uphill, and then back down.

Re: Grok 4.6

#493
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise.

So what you experience as a "near simultaneous" release is just their decision of when to peel off a release from their current set of in-training models, likely based on how they perceive market and regulatory conditions. They likely see a competitor release and then baseline what they should release based on that and it takes a month or two for them to package it up and push it out the door.

What I can imagine is that for some of the labs, they are being forced to publish models closer and closer to the frontier of what they have in training. Effectively, "falling behind" is your forward pipeline shrinking. Google ran out of forward pipeline. So far Anthropic and OpenAI didn't - but probably, one is shrinking.

Re: Grok 4.6

#494
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

Also, "criminal activity" doesn't have the same definition across jurisdictions. Seems like it would either be overzealous in its refusals or be easy to jailbreak by claiming a jurisdiction that is loose.

Re: Grok 4.6

#495

Earlier quoted context omitted.

> The alternative is Claude-style "safeguards" aka censorship Another obvious alternative is to just have the model do what you tell it to do, and then arrest people who use generic tools for crime instead of trying to make a kitchen knife that can't be used for stabbing someone.

[flagged]

"From a safe distance"

I.E. you haven't seen anything. You've just heard the same bullshit stories repeated ad naiseaum by haters.

I use X plenty every day. I've seen zero. Adult material right after Imagine was released sure, then even that was clamped down on.

Re: Grok 4.6

#496
post #477

Earlier quoted context omitted.

I think the comment you replied to was referring to the fact that when Twitter was taken over the entire Trust and Safety team was done away with. This has allowed child sexual abuse material to flourish on the platform.

The child abuse material problem was much worse before Twitter was taken over.

This. They ALLOWED it to exist. Now it's clamped down on where seen, personally I've zeen zero having used X every day since the liberation.

Re: Grok 4.6

#497
post #43

Tangental, but has anyone else noticed grok's voice mode got stupid and terse ~2 weeks ago? I've absolutely loved grok's voice mode since it came out (incredibly useful for brainstorming on walks and helping conceptualise and get the verbiage for expressing ideas) but it seems so have lost about 40 IQ points recently, and if the question is multi-part, it often answers just one part with no elaboration or explanation…

Yes, that was Grok's strongest point for me previously, and it's been basically unusable these past few weeks. I think there have been a few posts in the Grok subreddit (maybe on r/LoveGrok). It has a lot less personality which is a shame, but I'd take that if the answers themselves were good - but they lack information, have zero nuance, and repeat themselves pretty often too. Such a disappointing change.

Re: Grok 4.6

#498
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Do they have Fable level models? Or are they all saying “hey we’re dangerous too!” and hoping to get some token spend out of it?

Grok currently is comparable to Fable, above Opus, but several times cheaper. I rather use it and not see just a few prompts eat thru my quota

Re: Grok 4.6

#499
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

And you get it to run the answer by "What would Elon do?" Before outputting, so you get the best of both worlds :/

Re: Grok 4.6

#500
Having used Chatgpt, Claude for coding tasks till recently, am quite happy to be finally able to rely on the model I use when I also want unbiased output to be the top dog there too!
Post reply on HN