What version the intelligence vs cost graph is using? they didn't ran v4.2 to all models.
Artificial Analysis Intelligence Index v4.2
11–20 of 71 posts
Re: Artificial Analysis Intelligence Index v4.2
#12Re: Artificial Analysis Intelligence Index v4.2
#13Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…
Re: Artificial Analysis Intelligence Index v4.2
#14Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…
I don't agree. The main issue with their scoring/methodology is that the numbers make it seem like 5-6 models have little to no difference when in fact there is a significant difference between fable and opus and sol and astra for example. They are popular mainstream but most of their benchmarks are either not a representation of model strengths enough or they are not doing a good job of showcasing it properly. The f…
Re: Artificial Analysis Intelligence Index v4.2
#15This is really a great achievement: "Astra dominates the output token frontier" Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well. Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin. Edit: Just to make the point: Astra (max) has t…
This also makes it much harder to monitor its reasoning.
Re: Artificial Analysis Intelligence Index v4.2
#16This update really gives OpenAI a boost. Not saying there's anything inaccurate or untoward about that, but the timing is unfortunate. It would have looked better had it been done prior to the Fable 5.1 and GPT 6 releases. I guess AA would say that there's no perfect time to do these updates, given the rapid fire pace of releases!
Re: Artificial Analysis Intelligence Index v4.2
#17This is really a great achievement: "Astra dominates the output token frontier" Many labs used increased thinking to boost benchmark scores and performance. Most of the Chinese models were doing that for a while. Google and Anthropic as well. Not OpenAI. 5.6 already was much more token efficient than other models and Astra beats Sol in token efficiency by a wide margin. Edit: Just to make the point: Astra (max) has t…
I don't really understand why they keep doing this. Either run and report everyone at multiple effort levels, or run everyone at only one.
But alsi, token efficiency seems pretty artificial? For example tokenizers are different from model to model. The cost/perf Pareto frontier seems a lot more meaningful (and Astra does very well at that too, just to be clear. It seems to be a great model.)
Re: Artificial Analysis Intelligence Index v4.2
#18Imo the omniscience index they have has the highest correlation to actual usefulness of the models. https://artificialanalysis.ai/evaluations/omniscience > measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful be…
IMO, what really makes a model useful is its ability to process information within its context reliably and faithfully. I don’t care if it hallucinates George Washington’s favorite color, but I do care about it hallucinating the results of tool calls.
Re: Artificial Analysis Intelligence Index v4.2
#19They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
Completely discredits the index if it just gets modified to match social media vibes.
Re: Artificial Analysis Intelligence Index v4.2
#20How to view the previous version to compare?