Live data from Hacker News

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

artificialanalysis.ai

241–250 of 472 posts

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#241
post #47

Earlier quoted context omitted.

That's pretty surprising. Idk about Grok 4.6, but Grok 4.5 was clearly below Fable, Opus 5 and GPT 5.6 Sol.

It's a different type of model. In my admittedly judgemental observation, people that aren't the type to configure fully automated harnesses with good tools and skills and verifiers for their infrastructure and are way more interventionist in the way their agent works tend to like grok 4.5 more as the main agent. It's much faster and writes more simple and normal code that aligns a bit more with human written code. A…

Yeah [1] is really a thing and SlopCodeBench (https://www.scbench.ai/) kind of measures that.

You need to manually push models to clean up the slop every now and then otherwise it becomes chaotic. And every change with LLMs is always extra lines.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#242
post #31

Earlier quoted context omitted.

Probably can't advance frontier math yet, yeah. But please let us know other places you want to see Grok improve for future models!

What's the latest on Composer 3? Is it a sort of distilled Grok?

We plan to eventually have another model at that weight class, but right now trying to train the best possible model.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#243
post #47

Earlier quoted context omitted.

That's pretty surprising. Idk about Grok 4.6, but Grok 4.5 was clearly below Fable, Opus 5 and GPT 5.6 Sol.

It's much faster, so if you're not doing something cutting-edge, or you're doing the planning yourself and just using the LLM for implementation, the speed benefit outweighs the extra smarts of Fable/Opus5/Sol.

I've used all of the models extensively and Grok is only "faster" because it claims to be done minutes after you ask it to do something. It does not produce results anywhere near the Anthropic or OpenAI models, it just hacks a tiny piece of what you ask for and says "I'm done!". I also notice Musk-isms leaking through the model. Multiple times it's told me "this is not a roast" or "I'm not roasting your code". People who use this model because they align with Musk's ideology are doing us all a favor and weeding themselves out of the competition.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#244
post #187

SpaceXAI is the only frontier model company that had its own compute/date centres and soon chip making factory, I think they will pull ahead with cheaper tokens similar intelligence and better harness/tools. Grok build is 2-5x faster than Claude Code in my opinion.

>> I think they will pull ahead with cheaper tokens similar intelligence they just increased cache read from 0.30 to 0.50 - this has the biggest impact on agentic coding. Elon companies have the most expensive everything: xAI sub: $30 when other starts at $20, pro like sub for $300 where other charge $200. Expensive electric cars, powerwalls, solar roofs when competetive products/better are cheaper.

The cache reads are really insanely expensive.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#245
post #50

Earlier quoted context omitted.

The issue is that this is supposed to be a tech discussion forum and many of us don’t want to be repeatedly beaten over the head with other people’s politics. Plenty of other places to be political, we just wish this wasn’t one of them.

> many of us don’t want to be repeatedly beaten over the head with other people’s politics I agree completely. That's why I won't use Grok: its owner repeatedly beats us over the head with his politics, and I won't encourage it.

Huh. I feel that way with the other models - all the same. Eh, what can you do. Have fun out there.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#247
post #152

Earlier quoted context omitted.

Competition keeps service quality high and pricing low - even if you aren't using Grok, the mere existence of Grok keeps pricing for whatever provider you use lower and service faster and more reliable.

That's fair. But my point is from a business POV, why would SpaceX want to invests hundreds of billions of CapEx on a third or fourth frontier model, which cannot compete with chatGPT and claude on the high end, and getting squeezed by open weight models on the low end

They came here in three years. I think there’s a chance they will release the absolute top model soon. And I think they believe that as well. They have the compute. They have the money. They have the engineers.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#249

Earlier quoted context omitted.

> MADISON, Wis. (AP) — Billionaire Elon Musk likely broke Wisconsin law when he promised to hand out $1 million checks to voters in the 2025 state Supreme Court election, a bipartisan panel has found. > The Wisconsin Elections Commission last week referred two complaints to the Brown County district attorney’s office, which can choose to bring criminal charges over violating the state law against election bribery. Pr…

[flagged]

"Likely broke" is according to the Wisconsin Election Commission, as is attributed at the end of that very sentence.

> Encouraging people just to vote

And "Criminal Conspiracy" is just making plans with friends. Just because you can describe it in vague terms doesn't make it A-Okay.

> It's going to backfire hard for all the ad spend by MTV, Meta, and Google

It will not, because running ads is definitely not election bribery, whereas what Musk did likely is under Wisconsin law.

Re: Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

#250

Earlier quoted context omitted.

>their subscription now goes way further than OpenAI or Anthropic. Until it doesn't... Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter. And I am feeling a lot better about it now that I've finally got it working.

But still for US frontier you're paying 10-20x more per token compared to their limited subscriptions. For China frontier you'll be good though, and that might be the future anyway.

Grok is cheaper vs real Chinese frontier aka kimi. Sponsored or not.
Post reply on HN