Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

201–210 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#201
post #50
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

Replying to myself, because I just bumped into these: https://www.reddit.com/r/claude/comments/1vfvdgz/anthropic_l... https://www.reddit.com/r/ClaudeAI/comments/1vgpyni/my_opus_5...

Especially the second one seems exactly like my experience.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#203
post #85
post #41

Earlier quoted context omitted.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and keep your work targeted (you have to know what you want the thing to do!), clear your context, the models will do what you ask pretty reliably.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#204

Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol Source: http://pellmell.ai/leaderboard . This jumps around a lot based on the top throughput and latency of whatever provider happens to be b…

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

It takes 2 minutes to fix Opus 5

https://code.claude.com/docs/en/output-styles

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#205
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.

Works great with room to spare on my lenovo pgx too

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#206
post #168

Earlier quoted context omitted.

A smarter model would know how to communicate with you correctly, and not just throw jargon it has just invented at you without explaining it.

But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.

Effective jargon usage is understood by the target audience.

If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#207
post #38

Earlier quoted context omitted.

I use https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions. OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.

Is there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?

Opencode Go has Qwen 3.8 Max at $10/month

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#208
post #168

Earlier quoted context omitted.

A smarter model would know how to communicate with you correctly, and not just throw jargon it has just invented at you without explaining it.

But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.

Concise, jargon free or limited explanation is the opposite of dumbed down. It requires to most skill and understanding to do well. Opus 4.8/5, for whatever reason, are getting worse at this crucial skill.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#209
post #192
post #177

Earlier quoted context omitted.

yep matches my experience completely But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.

Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the d…

Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following:

"give me this again without jargon invented this session at high density

and with a couple (maybe more or less) simple useful ascii diagrams underneath each design"

The context is that I was discussing an experimental new idea for my video game review analysis product.

Designs 1,2, and 3 were horrible: the model even suggested a rejection after the word soup so it would have been pointless to waste my fleeting time on Earth reading it.

Otherwise, I generally really enjoyed using fable for bouncing ideas. It was an absolute joy to have this thing provide useful criticism, analyse sample data, and create prototypes so that I could elevate my understanding of the problem without stepping down from a pure intuition/design headspace.

But I don't consider the purely model written code usable for a feature this important. I'll probably scrap it entirely and start from scratch with newfound understanding.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#210
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
Post reply on HN