Live data from Hacker News

Artificial Analysis Intelligence Index v4.2

artificialanalysis.ai

51–60 of 71 posts

Re: Artificial Analysis Intelligence Index v4.2

#51
post #48

I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as the 2.1 results might suggest. The 4.0 results differentiate these models much more effectively. (2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)

DeepSWE has Gemini 3.8 Flash up really high, too.

Re: Artificial Analysis Intelligence Index v4.2

#52

Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…

I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The f…

I dont think you read my message. Muse and sol are nowhere near fable and Astra on the omniscience index

Re: Artificial Analysis Intelligence Index v4.2

#53

Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…

Hallucinations are very damaging to a model’s utility. But doesn’t the Omniscience Index focus on knowledge-based queries? To me, using LLMs for their memorized knowledge is very 2023 and suboptimal. IMO, what really makes a model useful is its ability to process information within its context reliably and faithfully. I don’t care if it hallucinates George Washington’s favorite color, but I do care about it hallucina…

Fair but I think those two behaviours are strongly correlated at least the index does represent my personal experience very well where fable is way better than opus opus is better than sol. And I haven't tried Astra yet but it having 44/43 is very interesting at least.

Re: Artificial Analysis Intelligence Index v4.2

#54

Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…

> A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses.

It's not useful because the benchmarks often measure the wrong thing. They're here yapping about AGI and yet the benchmarks treat it like a trained dog. Fetch this. 100 points.

Each "problem" in these benchmarks likely has more than 1 solution that can be considered correct and even should be graded in many ways. Yet we see in many benchmarks higher effort (or thinking levels) don't help because the benchmark penalizes for doing "more" than what the answers asks for. So what did you ask for?

In human school you often get marks on the process and not just the end result. Thinking tokens have been cut. All we group on is things like cost, turns and time but not the what else.

Re: Artificial Analysis Intelligence Index v4.2

#55
post #48

I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as the 2.1 results might suggest. The 4.0 results differentiate these models much more effectively. (2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)

DeepSWE has Gemini 3.8 Flash up really high, too.

It does and that one also feels kind of saturated for measuring the most advanced models - opus, Gemini, astra, sol, fable, glm, kimi all scoring within statistical noise of each other. (74 +-3% down to 69% +-5% for kimi).

It's still providing strong discrimination between weaker models.

Re: Artificial Analysis Intelligence Index v4.2

#57
post #30
post #6

Earlier quoted context omitted.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.

[deleted]

Re: Artificial Analysis Intelligence Index v4.2

#58

I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set? A glance at their new index shows that whatever they're measuring, it isn't useful. Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad. I haven't tried muse spark 1.3. B…

Because it's the best option currently available.

It's really easy to shit on AI benchmarks, but that noise is useless unless you're offering a solution or a better benchmark.

Re: Artificial Analysis Intelligence Index v4.2

#59
post #30
post #6

Earlier quoted context omitted.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.

1) open access is a rising trend in journal publishing. arxiv exists because of it but this also means that a lot of papers get 'published' before they are peer-reviewed which is often when significant issues are caught

2) there are definitely problems with journals and publishing but, like many similar absurdly reductive pronouncements, the argument that the current mode of scientific inquiry is bunk is both wrong and lacks nuance

the problem with modern publishing is that private equity is buying up publishers [0]. these publishers are then giving peer-reviewers no time and zero pay to do the necessary work of review [1] while also charging exorbitant rates for access. this is leading to worsening quality of the published research along with highly overburdened researchers who are stuck between shrinking funding [2] and their myriad other professional obligations

to just say that 'journals' are bunk is ignorant in a harmfully anti-empirical way. the process of empirical research and review is the entire reason why we see realworld results. foundations comprised of bullshit crumble fast but for some reason or another our economic system is highly driven to enshittifying everything it touches

to have defenders who claim AA is 'pushing the boundaries of the scientific method' sounds like the screeching refrain of anti-intellectual cargo cults, apeishly mimicking the features of rigor and methodology while avoiding any real accountability

[0] https://issues.org/how-academic-science-gave-its-soul-to-the...

[1] https://www.insidehighered.com/news/faculty/books-publishing...

[2] https://www.nature.com/articles/d41586-025-00754-4

Re: Artificial Analysis Intelligence Index v4.2

#60
post #6
post #5

They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repea…

[deleted]
Post reply on HN