Live data from Hacker News

Grok 4.6

x.ai

481–490 of 696 posts

Re: Grok 4.6

#482

Earlier quoted context omitted.

Then why is he left-handed?

Because the primary is right handed. I enjoy asking my grandkids why mirrors reflect left to right and not top to bottom.

They do "reflect top to bottom".

Hold a written word in front of your eyes to read it.

Now flip it to the mirror to read the reflection:

Did you flip it horizontally? Then it reflected left-to-right.

Did you flip it vertically? Then it reflected top-to-bottom.

Re: Grok 4.6

#483

Earlier quoted context omitted.

But in theory you can make an LLM A LOT faster than a human. You can also run massive amount of LLMs in parallel. There might be a limit to a normal LLM but not to theo everall system.

Humans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.

I think using different AGENTS.md can give the same model different perspectives on the same problem. For example a model with a well-tuned AGENTS.md by an expert mathematician approaching the same problem as the same model with a well-tuned AGENTS.md by an expert biologist can grind on the same problem from different perpectives.

It's worth a shot at least, as a microservices architect I have a bias that we aren't networking these enough, a single main agent session orchestrating multiple subagents is different from multiple main agent sessions with their own subagents coordinating with each other.

Re: Grok 4.6

#484

Earlier quoted context omitted.

But in theory you can make an LLM A LOT faster than a human. You can also run massive amount of LLMs in parallel. There might be a limit to a normal LLM but not to theo everall system.

Humans, however, are highly variable, which may produce really varied and interesting results if they work together. One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.

> 100 instances of the same LLM may converge on the same result as 10.

Not in the highly verifiable domains. There you can take it from say 80-90% maj@x to 99% pass@n. Math, some parts of programming and cybersec are examples of highly verifiable domains. (e.g. if you're searching for a linux LPE, that's expensive to search but easy/cheap to verify - just have a token in /root and have the model retrieve that token)

Re: Grok 4.6

#485

Earlier quoted context omitted.

> Remigration is a far-right concept referring to the ethnic cleansing[1] via mass deportation of non-white minority populations, especially immigrants and sometimes including native-born citizens, to their place of racial ancestry.[2] https://en.wikipedia.org/wiki/Remigration It’s right there at the top. One google search is all it takes. You didn’t even, for a second, think to familiarize yourself with the remigrat…

I reject your (and Wikipedia’s) theory of “remigration” arbitrarily defining something that is (not actually) happening in the US. People are being deported from the US based on their lack of legal presence in the country, which is what every country also does. Nobody is being deported because of their skin color. Also, nobody is jumping to the conclusion that you’re wrong about something that is not actually happeni…

No one is saying it is currently happening, you goof. Remigration is an idea proposed by identitarian assholes like https://en.wikipedia.org/wiki/Martin_Sellner, author of a book called Remigration, who Musk has promoted in tweets.

It is not currently happening, but the richest man in the world is among those actively working to make it happen.

Again, you types confidently misunderstand a discussion.

Re: Grok 4.6

#486
post #64
post #23

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations: 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? 2) Distillation - also implausible for the reason above. 3) Benchmark hacking. AI companies h…

> 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months? The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse. Mythos (Preview) was publicly announced in April 2026 [1] which me…

> misuse

Well that’s not really true; it covers completely legitimate use also.

Re: Grok 4.6

#487
post #174

Earlier quoted context omitted.

> * Do not provide assistance to users who are clearly trying to engage in criminal activity. I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. Kind of amusing that we made it as far as we did as a species not really being able to explain how the human brain does it's most amazing tricks and then we just replicated it while still not really understanding…

Mmm, quite. > I don't know what we want to call this, but in my opinion, having to convince your tools is not computer science. My vote is "machine psychology".

Robopsychology, of course.

Re: Grok 4.6

#488
post #477

Earlier quoted context omitted.

That has been the case for a while now: https://en.wikipedia.org/wiki/Ashcroft_v._Free_Speech_Coalit...

I think the comment you replied to was referring to the fact that when Twitter was taken over the entire Trust and Safety team was done away with. This has allowed child sexual abuse material to flourish on the platform.

The child abuse material problem was much worse before Twitter was taken over.

Re: Grok 4.6

#490

Earlier quoted context omitted.

Opus 5 is terrible. I'd even say it's a step backwards from 4.8. I'm getting high error rates from it, and then it catches the error, and then it sometimes errors the error fix (!). Just today I had to switch another agent to Fable with the instruction, "Please clean up the mess that Opus 5 made, thanks" The other day, Sol called Opus 5's handoff (a skill I have that is basically a compaction, but just written to a f…

Same here. Regularly reverting back to Opus 4.8 after 5.0 being terrible. Anthropic does this all the time (ruins their models for users) while they screw around with system prompts. Oh but it's for your own good of course! They know what's best for us all, if we would just give them a monopoly. I can't wait until OpenAI/Grok/Chinese models surpass them enough that their main character syndrome and smug doomerism no…

Opus 4.5 gang here :)

I've reverted enough times I just pin this version.

Post reply on HN