Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

71–80 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#71
I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

I have screenshots of both. The description above the chart is the same in boh cases:

> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

What happened? How can the scores change so much in a few seconds?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#72
post #4

Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.

Running a large model on rented GPU is still meaningfully more private than handing your chat logs over to FAGA

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#74
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

They JUST updated their methodology:

https://artificialanalysis.ai/methodology/intelligence-bench...

Edit to provide AA's article explaining it:

https://artificialanalysis.ai/articles/artificial-analysis-i...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#75
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

Same, they just updated it. Hacker news effect?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#77
post #41

Earlier quoted context omitted.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.

I do the same, and generally have good results, but it does stupid things with gusto.

I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#78
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

It's my daily driver. I like it and find it noticeably better than Opus 4.8.

After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.

My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#79
post #59
post #26

A couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56. I wasn't able to find an explanation from them. Anyone knows what happened?

A wire transfer happened.

The kind of distillation guaranteed to work.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#80
post #42
post #6

Earlier quoted context omitted.

It's not enough that it's better? Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don't want it around any more. People can fine-tune it for particular workloads.

> It's not enough that it's better? It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.

What status quo? Just look at Openrouter's rankings: https://openrouter.ai/rankings

Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.

Post reply on HN