Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

181–190 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#181

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

Is that what it feels like when the models get smarter than us?

No the models are just ass at communication without being directed.

Try asking them to make useful diagrams for some stuff in a codebase, out of the box without excessive hand holding they don't make good choices about what's worth communicating and how to do it.

You see this in their pointless frontend copy all the time too.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#182
post #64
post #58

Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!

> an annoying personality I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.

Claude 4.5/4.6 - absolutely agree. Fable 5? From my (limited) testing, also reasonable to interact with.

Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#183

Earlier quoted context omitted.

Damn I thought it was my extra instructions, I swear everything it writes is in some shorthand with direct references to variables that literally nobody could figure out unless you literally just wrote that code 5 minutes ago. I had it stop writing comments altogether cause it was always four lines of complete and utter nonsense, and it doesn't even obey that rule half the time. Despite doing an extensive back and fo…

Everything is load-bearing with 3 measured blockers.

Don’t forget the smoking guns! I think these new models have been reading too many Agatha Christie novels.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#184
post #74
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-i...

In that case they should clearly label that this is a new benchmark.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#185
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

I'm still skeptical of the smaller models after the talent exodus a few months ago.

At this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#186
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#187
post #119
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.

Same.

I could not get Opus 5 to do anything without losing a few years of my life from stress.

Fable has been okay but I am doing ML work and not allowed to use it which feels insane.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#188
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."

Yeah, I told it to save in its memory that I don't want to have any more word salad!

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#189
post #74
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-i...

Can someone please explain what changed, when it happened, and whether it was surreptitious?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#190
post #168

Earlier quoted context omitted.

Is that what it feels like when the models get smarter than us?

A smarter model would know how to communicate with you correctly, and not just throw jargon it has just invented at you without explaining it.

But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.
Post reply on HN