Live data from Hacker News

Grok 4.6

x.ai

81–90 of 696 posts

Re: Grok 4.6

#81

As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities. Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price. I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation ma…

Curious - what is the main issue you find polarizing with grok?

I'd start here:

https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a...

And here:

https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal

I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here; and we use Chinese models (*hosted by US providers) for context.

Re: Grok 4.6

#82
post #68
post #39

Earlier quoted context omitted.

Agreed, but my suspicion is tied to the timing. Catching up eventually is to be expected. Having similar jumps in capability ready at the same time is odd.

Maybe "readiness" is quite a flexible category? You're mid-training for your next model; a rival releases something; you clear the boards and release the model without completing the training run?

Touche, aborted training runs probably do happen often. Closed model providers have zero incentive to announce a new model with less-than-best benchmarks.

Re: Grok 4.6

#83

Earlier quoted context omitted.

Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great!

Is that still a thing? I assumed they would have done something about it by now.

Yeah, that got stopped I think.

Re: Grok 4.6

#84

It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with. Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it…

> It's crazy that I'd literally trust a Chinese AI company with my data

It's crazy how much Chinese = bad the media or US companies have washed into you. Why lump it together?

Like any place and any company there are good and bad 1s.

It's not the Wild West over there...

Re: Grok 4.6

#85
post #80
post #44

Earlier quoted context omitted.

Does not explain timing

keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.

Yeah that would make more sense, it's probably a tight community and word gets around when something starts working.

Re: Grok 4.6

#86
post #52

[flagged]

The last time grok made these statement, I tried using it for my workflows and it did not perform as good as opus or even sonnet.

My guess is that xai benchmaxxes a lot but fails in actual capacity to produce good models.

Re: Grok 4.6

#87
post #66

It's crazy that I'd literally trust a Chinese AI company with my data over anything Musk is involved with. Like, even if you don't care about (or even like) his politics and can look past how unlikable he comes off as, the damage he's done to his own reputation in this domain just makes using his products like this a no-go. He's literally so rich that he can get caught personally looking through chat sessions and it…

[flagged]

[flagged]

Re: Grok 4.6

#89
post #46

Earlier quoted context omitted.

I'm pretty sure both Anthropic and OpenAI haven't necessarily been secretive that they have internal models that are much more capable than commercially available ones. It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right.

The "one in the chamber" is another good candidate that could explain the timing.

I think this is the right one, iirc 5.6 came out quite soon after Opus 5 etc?
Post reply on HN