Live data from Hacker News

Grok 4.6

x.ai

681–690 of 696 posts

Re: Grok 4.6

#681

Earlier quoted context omitted.

I think using different AGENTS.md can give the same model different perspectives on the same problem. For example a model with a well-tuned AGENTS.md by an expert mathematician approaching the same problem as the same model with a well-tuned AGENTS.md by an expert biologist can grind on the same problem from different perpectives. It's worth a shot at least, as a microservices architect I have a bias that we aren't n…

Crucially, does it make capabilities infinitely scalable? My comment just said that models may have a hard cap, and maybe doing specific setups like yours can make reaching it easier, but making the 'team' 10x larger after that optimal point may bring few to no improvements. Although I'm also not sure about just how much better models can really get with this technique. Ultimately you're still getting the same model…

I agree it doesn't make the capabilities infinitely scalable, wasn't arguing with that point. It's just an experiment. I'm not talking about "you are an expert mathematician, go", I'm talking about an expert encoding their heuristics into the AGENTS.md base context. Routing the model's attention to very different aspects of the same problem in the early context.

FWIW I mean if I have an AGENTS.md that encodes my software heuristics (use an interface in situations like X, here's how we name variables, etc.) it generates far cleaner code than if I don't.

Edit- mostly pointing out that stacking 10 base models vs. 10 models with sufficiently different base context isn't necessarily the same attention routing. I suppose I was thinking about tasks that don't have a concrete single answer.

Re: Grok 4.6

#682

Earlier quoted context omitted.

As funny as Mechahitler was it was more of a Microsoft Tay moment with the chatbot parroting what Twitter’s users were telling him without guardrails or a safe system prompt. It had nothing to do with grok’s or Musks pro nazi views (or lack thereof)

Yes it did. Elon Musk wasn't happy that his own chatbot was to left, so they 'adjusted' grok so often until it became mechahitler.

But that’s not what actually happened.

Re: Grok 4.6

#683
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

My theory is that it all boils down to better data and longer post-training period. Cursor got curated data from the trillions reactions of real world developers in real jobs. xAI bought is and used it for its post-training and got Grok 4.5 . Longer post-training on the powerful Colossus cluster helped it get Grok 4.6 , although both versions use the same model with the same number of parameters. Thus, both must use…

Now GLM 5.3! And they explicitly confirms my theory: "Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them."

Re: Grok 4.6

#684

Earlier quoted context omitted.

Notice how in your own example, the anti-particle is not the particle, and Bob is not adastra22.

I did notice that these things are labels and not objects.

Good point.

And it would revolutionize our current understanding of physics to find out that p = p̄

Re: Grok 4.6

#685
post #559

Earlier quoted context omitted.

I thought that's just the legal definition in any sufficiently developed countries?

People have been prosecuted and convicted here in Sweden for Japanese hand drawn CSAM. I think it comes down to different ideas of why the law exists. If you believe removing access to pornographic material for this category means people will have a harder time becoming pedophiles, then that's how the Swedish law makes sense. If you believe pedophilia is a tragic disease that we can't treat and that synthetic pornogr…

The usual Japanese talking point is that CSAM laws exist for rights and wellbeing of kids, which are real human being. If it's used/defined for cases that are not plausibly harmful to kids, such as some adults doing some things nowhere near kids using none of real materials, then it's often argued that the law is abusively defined/used, and/or plain broken. That talking point, incidentally, should logically validate against texts of laws, though sometimes it may not with the intents between the lines like you've brought up.

But that logic usually align perfectly with the texts of laws. Those laws aren't supposed to be adult pedos beating and correction laws. They're not pedo soft landing laws either, they are supposed to be parts of younglings protection laws, globally. Technically that should be all that matters. So there's that.

Re: Grok 4.6

#686
post #221

Earlier quoted context omitted.

[flagged]

A trait I share with dictionary editors is a preference for linguistic descriptivism, so for me it's not a real problem that the common definition of "sex" and the scientific use are different. Unfortunately, reality doesn't care at all about the categories humans create, so there's always some exception like the following two no matter how you try to cut reality at the joints with word definitions. Even in humans, w…

A wall of text that agrees with the fact: human males cannot be pregnant. Intersex defects prove the point that human males cannot become pregnant.

Thanks for playing.

Re: Grok 4.6

#687

Earlier quoted context omitted.

[flagged]

Are you able to differentiate between sex and gender, or is any nuance too complicated for you? Maybe Grok can explain it!

Glad you agree biological males cannot be pregnant. I guess reality has a conservative bias.

>Maybe Grok can explain it!

Glad that Grok is outputting facts.

Re: Grok 4.6

#688
post #493
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

It's a combination of (1) and something you don't list: I think the frontier labs all have multiple generations of undisclosed models in continuous training. There is no "end point" when it's magically "ready". It's just getting better and better all the time. What they release with a name and a version number is just a marketing / branding exercise. So what you experience as a "near simultaneous" release is just the…

I've been hearing this myth since ChatGPT came out - that the labs have superAI that they're just slowly trickling out as competition forces them to.

Everyone in silicon valley has a cousin who's supposedly seen Anthropic's new unreleased model that changes everything forever. I distinctly recall sitting in a work meeting where a coworker was insisting GPT4 was AGI that anthropic was too scared to release. It's amazing to me that this playbook is still working at least somewhat on folks.

Re: Grok 4.6

#689
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

I'm sure Grok 4.6 is not Fable level. Benchmarks are almost useless. Having said that, Grok 4.6 (1.5T params) is without a doubt way smaller than Fable, maybe a Fable sized Grok would be Fable level?

> Benchmarks are almost useless

I both agree with this but also find most insistence that Model A is the "best" is social hype divorced from capabilities.

Re: Grok 4.6

#690
post #129

In terms of using experience, I found Grok 4.5 to be way more pleasant to use than GPT 5.6 Sol and Claude 4.8/5. It just gets to the point, and is super fast and concise, no yapping. That's how AI agents should be imo. None of the weird "Claude ipsum" jargon like "load-bearing" and "stale folklore" or GPT 5.6-isms like "focused regression" and "provenance".

What harness are you using?

These differences are prominent even in the same harness (cursor) in my experience. Opus is wordy to the point of exhaustion, GPT is better, Grok is very direct.
Post reply on HN