Grok 4.6
461–470 of 696 posts
Re: Grok 4.6
#462https://youtube.com/live/CjM6U7W7pk4
Here is the resulting page it built:
https://robss2020.github.io/frontier-brief/
Sorry that I didn't think of some larger project to build or something. It was kind of late.
Re: Grok 4.6
#463Re: Grok 4.6
#464Earlier quoted context omitted.
I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…
> I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible How could we really know how much "compute China has" in reality? Is it possible that whatever estimates people has come up with for both China and the US might not be 100% accurate?
Re: Grok 4.6
#465Re: Grok 4.6
#466Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline. I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they s…
(On mobile so can't search, but this was yesterday:)
> "Oracle was providing a staggering 22.6 percent of China's known A.I. computing power"
In Malaysia, etc.
Re: Grok 4.6
#467Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability. All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.
It (mythos) was first made public in April so it's not a surprise that others would catch up, though.
Re: Grok 4.6
#468Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…
The moral of the story? People work in parallel on the same goals, they build on best practice, or sometimes just need to see something is possible (reusable rockets). Having achievements cluster like this is normal and expected.
Re: Grok 4.6
#469Looks like the SpaceXAI api is adding a default system prompt to all requests. Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts """ You are Grok, a helpful and maximally truthful AI built by xAI. Your purpose is to answer questions accurately, be helpful, and seek truth above all else.…
Source for this? This seems like a crazy leak if it's their real system prompt. I find it hard to believe since I have tried system prompts like this and it doesn't work that well, just pollutes the user's context. A great test for any LLM is to ask its name - Mistral will respond with all kinds of stuff, sometimes other models' names, revealing that it has trained on other models. Grok doesn't though. It is "witty a…
Re: Grok 4.6
#470Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…