Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

251–260 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#251
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

Or are humans more of a bottleneck than before, because to improve on the most complex problems that demonstrate intelligence you need some way to verify that they are correct. If it's hard for humans to even know if something is correct, wouldn't that slow everything down and simply put limits on the scaling speed of models based on human verification?

So instead of relying heavily on human bottlenecks, you focus on agentic task verification since that's the low hanging fruit and verifiable at scale?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#252
post #181

Earlier quoted context omitted.

No the models are just ass at communication without being directed. Try asking them to make useful diagrams for some stuff in a codebase, out of the box without excessive hand holding they don't make good choices about what's worth communicating and how to do it. You see this in their pointless frontend copy all the time too.

Like any time you make them do any UI without strict directions they'll almost always add a label describing the feature somewhere. Ask for a calculator, and instructions for what the different buttons do might appear in the bottom out of nowhere for example. Same concept of "over-sharing" seems to prevalent in a bunch of domains when it comes to LLMs, sometimes more visible, sometimes less.

[dead]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#253
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

And I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem."

I suspect that the models are genuinely close and that certain experiences get felt across providers but are inconsistent enough to convince people one is superior to the other. I for one have tried Deepseek on and off since my co-founder is fond of it and I've stopped trying now because I never have a good experience.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#254
post #121

Earlier quoted context omitted.

I thought they are making a profit on API pricing? A quick Google shows somewhere between 50-70% margins on API inference.

API pricing is almost definitely profitable, but at this point I assume it's a small minority of their inference traffic compared to subscription usage, and unlikely to make up for the rest of their expenses on its own.

Why would you assume so when companies 150+ people can only use API pricing? My assumption is that more people use Claude at work than personally.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#255

Earlier quoted context omitted.

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

[deleted]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#256
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#257
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#258
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#259

Earlier quoted context omitted.

If the legal system declares the first thief’s theft not theft then all bets are off.

> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.

True. It's the courts that failed humanity. Or perhaps the shits that invented copyright to start

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#260
post #44

Earlier quoted context omitted.

Sent to solve one task, came back with half of it solved and 2 more problems.

What really enrages me is the amount of effort it puts into justifying weaseling out of work. (THAT'S MY JOB!) It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.

sounds like someone needs a local llm.
Post reply on HN