Live data from Hacker News

Grok 4.6

x.ai

381–390 of 696 posts

Re: Grok 4.6

#381
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Your implicit assumption seems to be that they didn’t start on this model until after Fable was released. They never stopped training though.

Re: Grok 4.6

#382

Earlier quoted context omitted.

You only mention math, coding and videogames. They already hire and pay people with research titles for creating and solving problems in their fields. And a lot of labs say that RL can help everywere and has plenty of way to go.

Correct. You can look at the AI tutor jobs in the job board of any of these companies.

Note that those jobs are miserable: https://nymag.com/intelligencer/article/white-collar-workers...

Re: Grok 4.6

#383
post #290

Earlier quoted context omitted.

Maybe because frontier labs buy the same RL tasks from task producer companies.

Who are these task producers? Are you saying that Anthropic, et al delegate the RL part to third party companies that do it for pretty much every other AI company as well?

There are companies that will pay you $$$ for technical challenges that stump frontier models. I’ve met these people. They make good money.

Re: Grok 4.6

#384
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Could it be that there's no magic formula, everybody uses the same known ideas, the same computation power, the same training data? if that's the case, we can imagine that models will be commoditized.

Re: Grok 4.6

#385
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

not suspicious at all. They are all doing the same scaling of test time, training data so getting similar results. anyone with access to capital can produce frotier model. hell you can just ask chatgpt how to create a fontier model. recipe is not a secret despite what these 'labs' pretend

Google is not able to currently produce a frontier model despite all the capital.

Re: Grok 4.6

#386
post #180

Earlier quoted context omitted.

If I was Chinese, I'd probably trust Grok more than a local AI company. Americans would probably trust the Chinese companies more. It's less about "who is more trustworthy", it's more about "who is more willing and able to affect me".

I think that underestimates how little the Chinese care about what Americans are doing. They're moving so fast that watching what the U.S. is doing would slow them down.

The Chinese do care what the Chinese government does, and are interested in minimizing what the government knows about them, and are well aware of the internet firewall the government operates and that any Chinese company will give the Chinese government whatever they ask for.

Re: Grok 4.6

#387

Earlier quoted context omitted.

They will get sharply better in tasks with verifiable domains... math and coding Gradually the labs will start engineering verifiable sandboxes for wider domains like videogames This strategy will hit a plateau in about 18 months and then we're back to diminishing returns and incremental progress along other dimensions (like accelerated inference using ASICs)

You only mention math, coding and videogames. They already hire and pay people with research titles for creating and solving problems in their fields. And a lot of labs say that RL can help everywere and has plenty of way to go.

RL can do behavior cloning, but really needs good simulations or verifiable environments to get to superhuman levels. That currently exists for math, coding, and a lot of videogames. Soon there will be good enough simulations for robotics.

There's a lot of domains where that simply isn't the case (like bio)

Re: Grok 4.6

#388
post #379

Earlier quoted context omitted.

We didn’t replicate the human brain. We built systems that can statistically approximate some of what the human brain might output in certain limited situations.

What do you think a human brain is…

This is like saying the person you see in the mirror is categorically a human being because both of you produce similar reflections of light rays

Re: Grok 4.6

#389
Does anyone know how the grok allowances compare to OpenAI / Anthropic for the monthly plans? I heard they're not generous, which means I never really bother testing Grok.

Re: Grok 4.6

#390
The model could be a legit, real life Jarvis sentient super brain genius and I would rather stay ignorant than give Elon any of my money.

I’m hoping all his enterprises burn to the ground. I’m glad there’s plenty of competition from China at far cheaper rates.

Post reply on HN