Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

171–180 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#171

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

Is that what it feels like when the models get smarter than us?

No, a smart model should also give a concise executive summary, "brevity is the soul of wit".

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#173
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

I'm still skeptical of the smaller models after the talent exodus a few months ago.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#174
post #74
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-i...

I have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#175

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

Is that what it feels like when the models get smarter than us?

"There is a view in some philosophical circles that anything that can be understood by people who have not studied philosophy is not profound enough to be worth saying. To the contrary, I suspect that whatever cannot be said clearly is probably not being thought clearly either."

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#176
post #168

Earlier quoted context omitted.

Is that what it feels like when the models get smarter than us?

A smarter model would know how to communicate with you correctly, and not just throw jargon it has just invented at you without explaining it.

s/model/engineer

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#177
post #50

Earlier quoted context omitted.

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard. With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or ma…

yep matches my experience completely

But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#178

Earlier quoted context omitted.

I find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation)

It's well known 35b is much faster (on any hardware) and quite a bit dumber

This really very much depends on how you are using it, I think. If you intend to leave it to solve long context problems and write whole prototypes, the 27B is going to be much better.

But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.

I would rather work that way, so I hope they do produce a small MoE model.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#179

Earlier quoted context omitted.

Is that what it feels like when the models get smarter than us?

No, a smart model should also give a concise executive summary, "brevity is the soul of wit".

So, in short, once models get smart enough they stop bothering telling us what they did.

Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#180
post #96

Earlier quoted context omitted.

"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."

This is the hard-won load-bearing quote.

Belt and braces all the way down
Post reply on HN