Live data from Hacker News

Artificial Analysis Intelligence Index v4.2

artificialanalysis.ai

21–30 of 71 posts

Re: Artificial Analysis Intelligence Index v4.2

#21
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

Hopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.

Re: Artificial Analysis Intelligence Index v4.2

#23
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

It's a benchmark for brand new technology and you admit the old score was silly but you think updating it is unscientific?

Re: Artificial Analysis Intelligence Index v4.2

#24
post #21
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

Hopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.

> but those are never discussed.

They are literally frequently discussed and it's why there are different benchmarks for different domains.

Re: Artificial Analysis Intelligence Index v4.2

#25
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

What evidence do you have that they tweaked the index to fit what they thought people expected?

Re: Artificial Analysis Intelligence Index v4.2

#26
post #8

This is really a great achievement: "Astra dominates the output token frontier" Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well. Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin. Edit: Just to make the point: Astra (max) has t…

Duh, that's just a benchmarking artifact. If you run one model at 4x recurrence compared to another, you get 4x thinking without increasing the tokens. So of course it's going to dominate the perf at given output token level. The truly sensible comparison is perf -vs- thinking flops (because each model might be a different unknown size, but labs are very secretive about what they're actually running under the hood) or perhaps cost (which can be misleading because of subsidies, but is at least practically relevant in the moment).

Re: Artificial Analysis Intelligence Index v4.2

#30
post #6
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.
Post reply on HN