Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

81–90 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#81

Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?

DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0].

[0]https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#83
post #44
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Sent to solve one task, came back with half of it solved and 2 more problems.

"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#84

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs that would drain my wallet - with claude code its just another Tuesday. Still love it.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#85
post #41
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#86
post #38
post #30

Earlier quoted context omitted.

How CLI are you guys using for qwen and kimi?

I use https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions. OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.

Is there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#87
post #16

Strange that the page https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?

According to those graphs, Grok 4.5 appears to be the most cost-effective model.

$0.05 per task, Intelligence Index score 52 -> GPT 5.6 Luna max

$0.36 per task, Intelligence Index score 56 -> Grok 4.5 high

$1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max

$0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh

$1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#89
post #4

Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.

Running a large model on rented GPU is still meaningfully more private than handing your chat logs over to FAGA

The acronym is new to me: Facebook, Anthropic, Google, openAi?
Post reply on HN