Live data from Hacker News

Grok 4.6

x.ai

241–250 of 696 posts

Re: Grok 4.6

#241
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Sometimes you just need to know that something is possible, not exactly how it is done.

Re: Grok 4.6

#242
post #39
post #29

Earlier quoted context omitted.

4) There's nothing terribly special about Anthropic. No moat.

Agreed, but my suspicion is tied to the timing. Catching up eventually is to be expected. Having similar jumps in capability ready at the same time is odd.

There's also a bit of selection bias going on here because we forget about labs that don't have a jump and just focus on the ones that do. Notably Google is definitely not having that capability jump.

Re: Grok 4.6

#243

Earlier quoted context omitted.

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

System prompts are more like suggestions than hard constraints.

I don't understand why they don't look for large substring matches for the system prompt before returning the response. Trivial calculation compared to a system prompt instruction asking the model not to do it

Re: Grok 4.6

#244

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

The US govt trusts SpaceXAI for defense and high security missions. The idea they are lying about contracted AI services is absurd. They're also a public company which beings even more oversight than openai / anthropic.

[deleted]

Re: Grok 4.6

#245

Earlier quoted context omitted.

Curious - what is the main issue you find polarizing with grok?

Not the person you are responding to, but the fact that Grok is being used to generate a ton of CSAM and pornographic deepfakes isn't great!

It’s pretty telling that almost all of the bullet points in the system prompt that was posted for Grok have to do with preventing criminality and CSAM generation. No other provider has this same issue at that scale.

The first-order-thinking reaction is “oh cool, look how they don’t want it to happen” but the second-order reaction is “why does this company have such a problem when others don’t?” It’s their own tactics. If you want the “good” of 4chan-like behavior, turns out you get the bad too.

Re: Grok 4.6

#246

Earlier quoted context omitted.

I'd start here: https://en.wikipedia.org/wiki/Grok_(chatbot)#Controversies_a... And here: https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal I think polarizing is a generous way of describing the problems. My organization has outright banned Grok, because we don't trust SpaceX to hold up to contractual agreements vis-a-vis data-privacy/training. That's the level of reputational damage we're talking about here…

Can someone help me understand the deep fake controversy? That's like making photoshop illegal.

A lot of AI users are profoundly stupid and intently malicious. That changes perception of the tool... IMO it's because generative AI data is inherently toxic and contains elements that incite primal rage, but that's just my gut theory.

Re: Grok 4.6

#247
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I'm sure the SF AI scene leaks like a sieve, and companies have a pretty good idea what each other is working on.

Re: Grok 4.6

#248
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Maybe compute is the real moat (chinese possibly skip around it with distillation), xai is buildouts have been insanely fast (colossus 1 - 100,000 H100 GPUs brought online in 122 days lol) so maybe that explains them catching up asked grok to give a compute estimate for each: - SpaceX / xAI: ~1.4 GW (owned Colossus clusters) - OpenAI: ~2–3 GW (mostly rented/cloud) - Anthropic: ~1.5–2.5 GW (multi-cloud + xAI lease) ch…

> Maybe compute is the real moat (chinese possibly skip around it with distillation)

Makes no sense. At this point, all Western AI companies also engage in distillation. If distillation were such magic, they'd be insane not to.

Re: Grok 4.6

#249
post #57

Earlier quoted context omitted.

Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers. So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.

GPUs might explain the remarkably concurrent timing. Data access doesn't really explain it unless all labs simultaneously got access to some treasure trove of data.

> Data access doesn't really explain it unless all labs simultaneously got access to some treasure trove of data.

They have data from their competitors model outputs. It is very hard to serve an LLM without also exposing how it works.

Re: Grok 4.6

#250
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Maybe compute is the real moat (chinese possibly skip around it with distillation), xai is buildouts have been insanely fast (colossus 1 - 100,000 H100 GPUs brought online in 122 days lol) so maybe that explains them catching up asked grok to give a compute estimate for each: - SpaceX / xAI: ~1.4 GW (owned Colossus clusters) - OpenAI: ~2–3 GW (mostly rented/cloud) - Anthropic: ~1.5–2.5 GW (multi-cloud + xAI lease) ch…

we are in the process of transitioning from hype to commodity with llm tokens, moats are typically at the top of the stack or in the data warehouse
Post reply on HN