Live data from Hacker News

Artificial Analysis Intelligence Index v4.2

artificialanalysis.ai

31–40 of 71 posts

Re: Artificial Analysis Intelligence Index v4.2

#31
I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set?

A glance at their new index shows that whatever they're measuring, it isn't useful.

Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.

I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.

Video game journalism vibes all over this.

Re: Artificial Analysis Intelligence Index v4.2

#32
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

Maybe there is truth in it?

I asked Astra to make some changes, it built crazy overengineered code, switched back to Sol, said code is too complicated, rewrite it, and it made lot cleaner code.

I feel some models could chaise complicated benchmarks too much in expense of simpler tasks quality.

Re: Artificial Analysis Intelligence Index v4.2

#33
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

Well, do you have any ideas about how to make it scientific? And they clearly needed to do something based on their unnatural ranking

Re: Artificial Analysis Intelligence Index v4.2

#34
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

Maybe there is truth in it? I asked Astra to make some changes, it built crazy overengineered code, switched back to Sol, said code is too complicated, rewrite it, and it made lot cleaner code. I feel some models could chaise complicated benchmarks too much in expense of simpler tasks quality.

Sol frequently over-engineers code.

Re: Artificial Analysis Intelligence Index v4.2

#35
IMO one of the biggest losses of the OpenAI/Cursor breakup will be the loss of OAI models on CursorBench [1]. Their bench has always been one that most-fit my mental model of how good each of these models are. I find AA’s Intelligence index to often be out of alignment with my own subjective evals.

[1] https://cursor.com/evals

Re: Artificial Analysis Intelligence Index v4.2

#36

I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set? A glance at their new index shows that whatever they're measuring, it isn't useful. Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad. I haven't tried muse spark 1.3. B…

What other benchmarks do you recommend that are more accurate?

Re: Artificial Analysis Intelligence Index v4.2

#37
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

> Astra is way better than Sol

This needs to be evaluated per task, jagged frontier yadda yadda. I would be not at all surprised if sol was better at some things than Astra just like people still use opus 4.6 and for good reasons.

Re: Artificial Analysis Intelligence Index v4.2

#38
post #26
post #8

This is really a great achievement: "Astra dominates the output token frontier" Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well. Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin. Edit: Just to make the point: Astra (max) has t…

Duh, that's just a benchmarking artifact. If you run one model at 4x recurrence compared to another, you get 4x thinking without increasing the tokens. So of course it's going to dominate the perf at given output token level. The truly sensible comparison is perf -vs- thinking flops (because each model might be a different unknown size, but labs are very secretive about what they're actually running under the hood) o…

But the end user doesn't care about flops for closed models. All they care about is how much it ends up costing them.

Re: Artificial Analysis Intelligence Index v4.2

#39

Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…

I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The f…

Also ArtificialAnalysis drop older foundational models to make room for new ones.

If you customize the filter to add Gemini 3.1 Pro you'll see it ranks 5th in AA-Omniscince Index. Yet by default you wont see Gemini 3.1 Pro.

I find this to be very unhelpful and confusing.

Re: Artificial Analysis Intelligence Index v4.2

#40
post #30
post #6

Earlier quoted context omitted.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.

It's not about journals. If they want to be 100 trustworthy they should release end to end reproducible pipelines for the whole process with everything but the private datasets included, and all design decisions fully documented. And let third party labs audit to confirm that the holdout questions are equal difficulty and similar task types to the public ones.
Post reply on HN