Live data from Hacker News

Artificial Analysis Intelligence Index v4.2

artificialanalysis.ai

41–50 of 71 posts

Re: Artificial Analysis Intelligence Index v4.2

#41

Earlier quoted context omitted.

I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The f…

Is there something better over there you'd recommend?

The Epoch Capabilities Index uses an Elo-based aggregation method that dynamically adjusts for benchmark difficulty and they put error bars on their scores, both of which put them miles ahead of Artificial Analysis: https://epoch.ai/eci?view=graph&tab=leaderboard

Re: Artificial Analysis Intelligence Index v4.2

#42
While I like this index, calling in "Intelligence" might be confusing - it is a mix of coding and knowledge.

Compare and contrast with ARC-AGI, BabaIsBench (https://quesma.com/benchmarks/babaisbench/), or MazeBench (https://mazebench.com/blog?post=introducing-mazebench).

In particular, in one Baba Is Bench post (https://quesma.com/blog/baba-is-aug-2026/), while quoting a Pareto frontier chart from AA, I noted:

> Intelligence Index vs. Cost per Intelligence Index Task from Artificial Analysis. Note that it is based on score of benchmarks like Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond - not necessarily fluid intelligence like in abstract puzzle games of ARC-AGI-3 or Baba is You.

Re: Artificial Analysis Intelligence Index v4.2

#43
post #21

Earlier quoted context omitted.

Hopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.

> but those are never discussed. They are literally frequently discussed and it's why there are different benchmarks for different domains.

> different benchmarks for different domains.

Labs know the only thing the public even discusses on model releases are benchmarks, so they devote a majority of training on just benchmaxxing. It’s marketing.

Muse Spark looks great in benchmarks. Everyone I know who has tried it (Rust & C++ projects) has determined it’s a resounding “meh”. That doesn’t mean it’s not a great tool for frontend devs, I wouldn’t know.

Re: Artificial Analysis Intelligence Index v4.2

#44
I wonder if Artificial Analysis could be influenced by certain model companies. Looking at the changelog [0], they updated a few times after new models appeared, so US models progression was much higher than that of other vendors (like when they updated the algorithm after Kimi K3 versus Opus 4.7, so Kimi dropped in the rankings). Or maybe thats just coincidence.

[0] https://artificialanalysis.ai/changelog

Re: Artificial Analysis Intelligence Index v4.2

#45
post #2

Why have they not included ARC-Agi-3 on their index? Clearly that would move things around.

Because playing a video game isn’t relevant to which AI people might want to use.

"Video game" is the medium, the challenge and test is figuring out how to win; when given no instructions, and specifically designed to be private.

ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.

Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.

Re: Artificial Analysis Intelligence Index v4.2

#46

I wonder if Artificial Analysis could be influenced by certain model companies. Looking at the changelog [0], they updated a few times after new models appeared, so US models progression was much higher than that of other vendors (like when they updated the algorithm after Kimi K3 versus Opus 4.7, so Kimi dropped in the rankings). Or maybe thats just coincidence. [0] https://artificialanalysis.ai/changelog

That would be the fastest way for them to completely torch their company.

The only thing they are selling and why people look at them is trust that they do honest evaluations.

Re: Artificial Analysis Intelligence Index v4.2

#47

I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set? A glance at their new index shows that whatever they're measuring, it isn't useful. Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad. I haven't tried muse spark 1.3. B…

As the saying goes, Artificial Analysis is the worst benchmarking company, except for all the others.

Re: Artificial Analysis Intelligence Index v4.2

#48
I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as the 2.1 results might suggest. The 4.0 results differentiate these models much more effectively.

(2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)

Re: Artificial Analysis Intelligence Index v4.2

#49
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

> Astra is way better than Sol

For knowledge type questions is that necessarily true? In the past we've seen things like Google's models degrading on general knowledge after the preview releases while improving on code/tool use, presumably due to catastrophic forgetting from the additional training. Their preview would be free, get lots of agentic use from users, then additional training on that and probably additional automated RL.

However Astra is on an entirely new base model so I also wouldn't expect it to be worse.

Re: Artificial Analysis Intelligence Index v4.2

#50
post #36

I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set? A glance at their new index shows that whatever they're measuring, it isn't useful. Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad. I haven't tried muse spark 1.3. B…

What other benchmarks do you recommend that are more accurate?

Give a task you have to 3 different models and see what actually works for you. There are no good benchmarks.
Post reply on HN