Live data from Hacker News

Artificial Analysis Intelligence Index v4.2

artificialanalysis.ai

61–70 of 71 posts

Re: Artificial Analysis Intelligence Index v4.2

#61
post #6
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

> but where's the outcome dataset justifying this? how did they get that probability? what was the specific methodology of the tests? what variables did they account for?

It's just high school statistics: https://en.wikipedia.org/wiki/Confidence_interval

It's not a probability. It's essentially an assertion that if the test were run 100 times, the result would be within the interval 95 times.

Re: Artificial Analysis Intelligence Index v4.2

#63
post #26

Earlier quoted context omitted.

Duh, that's just a benchmarking artifact. If you run one model at 4x recurrence compared to another, you get 4x thinking without increasing the tokens. So of course it's going to dominate the perf at given output token level. The truly sensible comparison is perf -vs- thinking flops (because each model might be a different unknown size, but labs are very secretive about what they're actually running under the hood) o…

But the end user doesn't care about flops for closed models. All they care about is how much it ends up costing them.

Which is why I said that cost might be a better metric than output tokens. But even that is somewhat misleading -- because there isn't really a single price -- there are a variety of offers / deals / subsidies, including subscription plans.

But if we're actually talking benchmarking intelligence -- as evidenced by "Artificial Analysis Intelligence Index" and the rest of the reporting -- then comparing systems with different amounts of recurrence on output tokens is flawed.

Re: Artificial Analysis Intelligence Index v4.2

#64

Earlier quoted context omitted.

I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The f…

Also ArtificialAnalysis drop older foundational models to make room for new ones. If you customize the filter to add Gemini 3.1 Pro you'll see it ranks 5th in AA-Omniscince Index. Yet by default you wont see Gemini 3.1 Pro. I find this to be very unhelpful and confusing.

But it is actually a great model it e.g. got the carwash question right from 9 months ago. While openais models all struggled.

Re: Artificial Analysis Intelligence Index v4.2

#66

This update really gives OpenAI a boost. Not saying there's anything inaccurate or untoward about that, but the timing is unfortunate. It would have looked better had it been done prior to the Fable 5.1 and GPT 6 releases. I guess AA would say that there's no perfect time to do these updates, given the rapid fire pace of releases!

The timing is related to the fact that their benchmark was saying it was the same as Sol, and below Opus 5, when anecdotal reports and other benchmarks strongly disagree. It looked bad for them for their benchmark to disagree with people's lived experience so hard.

Kind of reminds me when GPU benchmarks used to game the drivers to maximize the FPS. If the benchmark slightly changes the camera view is that cheating? Or is it calling out the cheaters?

Re: Artificial Analysis Intelligence Index v4.2

#68
post #59
post #30

Earlier quoted context omitted.

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.

1) open access is a rising trend in journal publishing. arxiv exists because of it but this also means that a lot of papers get 'published' before they are peer-reviewed which is often when significant issues are caught 2) there are definitely problems with journals and publishing but, like many similar absurdly reductive pronouncements, the argument that the current mode of scientific inquiry is bunk is both wrong a…

Yes, journals are harmful to freedom of knowladge which might seem ignorant especially to people coming from academia and mainstream who doesn't know better and can't imagine anything else besided wallen gardens they grew up with.

Arxiv is dominated by CS [1] and nothing really changed over the years [2]. Saying about pushing the boundaries of scientific methods it was another rock into journals and the status quo on how it's done like p-hacking and other data manipulations. Being in open with community notes would would have a difference.

Empirical research can exist without journals, science can be for masses and journals are antithetical for sharing making the knowledge exclusive.

[1] https://info.arxiv.org/about/reports/submission_category_by_... [2] https://en.wikipedia.org/wiki/ArXiv?useskin=vector#/media/Fi...

Re: Artificial Analysis Intelligence Index v4.2

#70

I wonder if Artificial Analysis could be influenced by certain model companies. Looking at the changelog [0], they updated a few times after new models appeared, so US models progression was much higher than that of other vendors (like when they updated the algorithm after Kimi K3 versus Opus 4.7, so Kimi dropped in the rankings). Or maybe thats just coincidence. [0] https://artificialanalysis.ai/changelog

That would be the fastest way for them to completely torch their company. The only thing they are selling and why people look at them is trust that they do honest evaluations.

People still buy into "research" from Gartner and Forrester despite it being an open secret for decades that they are pay-to-play.

These kinds of "research" companies are there to validate people's preconceived ideas and purchasing decisions rather than being genuinely unbiased. And they are very useful for that, both for consumers and for marketers.

Post reply on HN