Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

31–40 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#32

I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.

How are you running it on a Strix Halo? The weights aren't out yet, are they?

I interpret @syntaxing as meaning they are looking forward to running Qwen3.8-27B, but are frustrated by prefill times with other models, such as Qwen3.6-27B.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#33
post #23
post #15

Earlier quoted context omitted.

Indeed, and this Qwen 3.8 max specific page: https://artificialanalysis.ai/models/qwen3-8-max Doesn't have the claim either. Clickbait?

This page has it, scroll to "Intelligence" header (not the highlights one, but second on the page / with black square) and click "Agentic Index"

So the original link should be: https://artificialanalysis.ai/models/qwen3-8-max?intelligenc...

Even then, this seems a much more marginal win than the headline suggested to me.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#35

I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.

How are you running it on a Strix Halo? The weights aren't out yet, are they?

I meant Qwen3.6. Unsloth supposedly has early preview of the model and the VRAM requirement is the same so most people expect similar model size and type.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#36

Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?

Does "artificial analysis" mean what it says? Dubious.

But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me.

A lingering sense of quality degradation when going deep remains.

But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well.

And on-prem capability is simply irreplaceable.

Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#37
post #11

Earlier quoted context omitted.

For one thing, providers of open models can't arbitrarily increase their prices without facing competition.

But given the extremely low cost of switching, why wouldn't you use the cheaper one if they're comparable?

Speed and reliability.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#38
post #30
post #5

I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth. I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A vers…

How CLI are you guys using for qwen and kimi?

I use https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions.

OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#39

Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?

Almost surely. Apple is extremely well positioned to take advantage of this over the next decade.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#40
Hopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
Post reply on HN