Live data from Hacker News

Grok 4.6

x.ai

21–30 of 696 posts

Re: Grok 4.6

#22

Thats actually a lot more impressive than I thought. At least on paper

But has it hacked anybody yet? Feels like xAi is behind on the hot new benchmarking meta.

Didn't need to! The harness just uploads your repository to their blob storage directly. Cheaper than asking the LLM to do it

Re: Grok 4.6

#23
Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations:

1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

2) Distillation - also implausible for the reason above.

3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.

Other reasons?

Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

Re: Grok 4.6

#25
>Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass.

As a designer, I'm always hesitant to believe these statements until there's independent comparisons between the old & new model, as well as comparisons to human made flows. Design can be so subjective that blanket statements like this seem almost useless.

Re: Grok 4.6

#26
As polarizing as grok is, it was basically inevitable for it to start being a real competitor given how much investment SpaceX made into its own inference capabilities.

Seems if you are okay with it, there's no reason to use anything but the highest effort levels of some other frontier models for the price.

I think Grok provides healthy competition to the other labs, though I do think they bank on groks reputation making it less appealing to many.

Re: Grok 4.6

#29
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

4) There's nothing terribly special about Anthropic. No moat.

Re: Grok 4.6

#30
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers. So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.
Post reply on HN