Live data from Hacker News

Grok 4.6

x.ai

251–260 of 696 posts

Re: Grok 4.6

#251

Earlier quoted context omitted.

Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers. So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.

Yea, this reads as LLMs are a pretty obvious technology to develop(for the highly intelligent researchers who are there). Also there's probably a lot of actual divergence in model capabilities and skills that concealed by the fairly narrow set of tests we run them against nowadays. Like wasn't Grok 4.20 super targeted at non-coding tasks.

Why is everyone ignoring the pattern that has existed since training models became a thing? At first it sucks. Then it's better than humans. Just by using it you generate training data that makes it better over time.

Re: Grok 4.6

#252
post #29
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

4) There's nothing terribly special about Anthropic. No moat.

brand is their power, they'd be wise to not wreck it with dumb moves or PR statements (they already have some)

Re: Grok 4.6

#253

Earlier quoted context omitted.

> Annoyingly, the line about not mentioning these guidelines is superseding any instructions in the system prompt, causing the model to often refuse discussion regarding system prompts If the prompt guidance is causing the model to be so paranoid about leaking the system prompt... how do we already have it ?

System prompts are more like suggestions than hard constraints.

i beg to differ, in an ideal world a system possibly is a binding law and high end models are starting to be really aligned to the exact system prompt. The instructions must be simple to follow, if you start doing complex rules it'll call apart, but I'll usually follow the stringer interpretation.

Re: Grok 4.6

#254
post #159
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab. But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could mu…

But aren't today's frontier models already "fully universal"? To use your Turing machine analogy, I think we're past the calculator stage.

Re: Grok 4.6

#255
post #82

Earlier quoted context omitted.

Touche, aborted training runs probably do happen often. Closed model providers have zero incentive to announce a new model with less-than-best benchmarks.

I don’t think the runs need to be aborted… you can just release a mid-training checkpoint!

you'd be crazy to not be taking snapshots on the regular, many good reasons besides failures

Re: Grok 4.6

#256
post #159
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab. But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could mu…

You're take basically lines up with Francois Chollet: https://arxiv.org/abs/1911.01547

intelligence is more like polishing a ball smooth than growing the ball to infinity.

For many tasks, it will be smooth enough.

Re: Grok 4.6

#257
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

Why would you release a model if you are the current frontrunner? Only when a competitor pulls ahead, or comes close enough to actually get traffic, you prepare a new release.

Re: Grok 4.6

#258
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I'm sure the SF AI scene leaks like a sieve, and companies have a pretty good idea what each other is working on.

Okay so everyone is blaming diffusion or spying or whatever but we all use all of the models on our various projects in aggregate and they get to all read the code each other is generating. I do this with research tasks and local random stuff too.

So why do people have this idea in their heads that it's all some sorta secret sauce they are taking from each other?

Re: Grok 4.6

#259

Earlier quoted context omitted.

[flagged]

Not unless you're here illegally. And it has nothing to do with skin color. Just the basic fact that a country not in control of its borders ceases to be a country.

> Remigration is a far-right concept referring to the ethnic cleansing[1] via mass deportation of non-white minority populations, especially immigrants and sometimes including native-born citizens, to their place of racial ancestry.[2]

https://en.wikipedia.org/wiki/Remigration

It’s right there at the top. One google search is all it takes. You didn’t even, for a second, think to familiarize yourself with the remigration concept. You jumped immediately to me being wrong, even though I was discussing something you were ignorant of. That’s embarrassing.

Re: Grok 4.6

#260
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

That's exactly what Anthropic said was going to happen! Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.

They will get sharply better in tasks with verifiable domains... math and coding

Gradually the labs will start engineering verifiable sandboxes for wider domains like videogames

This strategy will hit a plateau in about 18 months and then we're back to diminishing returns and incremental progress along other dimensions (like accelerated inference using ASICs)

Post reply on HN