Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

291–300 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#291
post #85
post #41

Earlier quoted context omitted.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

At this stage in the game almost none of the comments or articles on HN can be trusted, if you know what I mean...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#292
post #290

Earlier quoted context omitted.

If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.

- load bearing -

seams, boundaries, envelopes, etc

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#293

Earlier quoted context omitted.

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.

Pareto optimal dominant vs a human for the same task, not an unreasonable framing but that assumes that it can actually do the task, which the op was arguing it couldn’t at all. Which, I suppose you could model as the utility of task completion % as being non linear. I have heard many people argue that the nature of work is messy and complicated and many things they do could not easily be emulated or automated. I do wonder how many of those activities are actually something that are connected to a companies ability to generate revenue or are just the messy interactions between people.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#294

Earlier quoted context omitted.

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

That’s definitely my impression of deepseek 0731 after a fair bit of use via ds4, it sounds like Claude.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#295

Earlier quoted context omitted.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.

[deleted]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#296

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

> So why would I want to switch to even worse model? Even if it's just slightly worse?

Self-hosting is the biggest reason.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#297

Earlier quoted context omitted.

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.

Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#298

Earlier quoted context omitted.

> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.

As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.

> They didn’t accept guilt for incorporating all of human output into their model without consent.

Because that use case is actually permitted by law.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#299
post #248

Earlier quoted context omitted.

> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.

> But they didn't find it. The Big LLM provider accepted guilt and paid a fine. That's not how it works. You have to give it back. Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?

[dead]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#300
post #96

Earlier quoted context omitted.

"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."

This is the hard-won load-bearing quote.

The shape of this problem is very heavy
Post reply on HN