Live data from Hacker News

Grok 4.6

x.ai

461–470 of 696 posts

Re: Grok 4.6

#461
The back and forth of llm's one up-ing each other is not worth the effort to keep switching harnesses/UI and established setup/workflow. Codex/Sol works well, i cannot imagine this to be a quantum leap in cost/efficiency/intelligence balance to spend the effort to make the switch a no brainer.

Re: Grok 4.6

#464
post #326

Earlier quoted context omitted.

I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…

> I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible How could we really know how much "compute China has" in reality? Is it possible that whatever estimates people has come up with for both China and the US might not be 100% accurate?

I'm not an expert but I think this sort of thing is relatively traceable for two reasons. One, datacenters are difficult to conceal. Two, the supply chains for many of the relevant materials are difficult to conceal. Some of those supply chains still require western components, I believe, so if you know how much of X component was sent to china, you know how much compute they have.

Re: Grok 4.6

#466
post #326
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…

There have been many reports that they're training in other countries.

(On mobile so can't search, but this was yesterday:)

> "Oracle was providing a staggering 22.6 percent of China's known A.I. computing power"

In Malaysia, etc.

https://thezvi.substack.com/p/ai-180-no-longer-in-charge

Re: Grok 4.6

#467
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.

Are you just guessing this? Using it, it's was clearly a jump in intelligence over previous models.

It (mythos) was first made public in April so it's not a surprise that others would catch up, though.

Re: Grok 4.6

#468
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

To the best of our recorded knowledge, nobody ran a 4-minute mile in the five millennia prior to Roger Bannister in May 1954[0], but more than 2,000 people have met or exceeded this achievement since. In fact, his record stood only briefly, being bested the following month by John Landy.

The moral of the story? People work in parallel on the same goals, they build on best practice, or sometimes just need to see something is possible (reusable rockets). Having achievements cluster like this is normal and expected.

[0] https://en.wikipedia.org/wiki/Four-minute_mile

Re: Grok 4.6

#469
post #164

Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…

Source for this? This seems like a crazy leak if it's their real system prompt. I find it hard to believe since I have tried system prompts like this and it doesn't work that well, just pollutes the user's context. A great test for any LLM is to ask its name - Mistral will respond with all kinds of stuff, sometimes other models' names, revealing that it has trained on other models. Grok doesn't though. It is "witty a…

Crazy? System prompt leaks are old news with dozens of trix to do it

Re: Grok 4.6

#470
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

There is a herd of companies all running a race. The technology is known. They all have roughly the same resources. It’s not unexpected that they have similar cycle times for model development and that those models will be of roughly the same quality. Then layer in corporate PR demands and you see all these models landing within weeks, sometimes days, of each other to keep the model developer’s name associated with “frontier” development.
Post reply on HN