Earlier quoted context omitted.
All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…
Is that what it feels like when the models get smarter than us?
Qwen3.8 Max now ranked as the best overall model by agentic index
171–180 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#172Well, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#173China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#174I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…
They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-i...
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#175Earlier quoted context omitted.
All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…
Is that what it feels like when the models get smarter than us?
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#176Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#177Earlier quoted context omitted.
Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.
I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard. With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or ma…
But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#178Earlier quoted context omitted.
I find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation)
It's well known 35b is much faster (on any hardware) and quite a bit dumber
But if you are sort of pair-programming with the model, the speed obviously matters and I think then the 35B is acceptably smart, and when it's wrong it'll be wrong much more quickly. It seems very good on SQL and PHP, and I assume on typical JS and Python.
I would rather work that way, so I hope they do produce a small MoE model.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#179Earlier quoted context omitted.
Is that what it feels like when the models get smarter than us?
No, a smart model should also give a concise executive summary, "brevity is the soul of wit".
Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#180Earlier quoted context omitted.
"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."
This is the hard-won load-bearing quote.