Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

321–330 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#321

Earlier quoted context omitted.

Is there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?

My first web search turned up this as the top result https://www.alibabacloud.com/help/en/model-studio/coding-pla...

Currently only these models are available: qwen3.7-plus (vision), qwen3.6-plus (vision), kimi-k2.5 (vision), glm-5, and MiniMax-M2.5

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#322

Earlier quoted context omitted.

Effective jargon usage is understood by the target audience. If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.

Or you're below its level.

That's a loadbearing, heavy shaped, second take worthy statement

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#323

Earlier quoted context omitted.

Fable has spoiled us all.

Not so sure, I'm sure Opus 5 is just shit.

Eh it's better than 4.8 in terms of what it can get done on a good day, it's just far more taxing to get it there.

Like the Fable ban stunt, I wouldn't put it pass Anthropic to kneecap Opus deliberately to drive more people to their more expensive option.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#324

Earlier quoted context omitted.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

question's phrasing made your judgement obvious

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#326

Earlier quoted context omitted.

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

I think it's more likely to be the effect of synchronization of launches, and the fact that models that do not challenge SOTA in some way do not get launched (think Gemini Pro delays), launched quietly or do not get any attention.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#327

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

And I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem." I suspect that…

That may be true.

I switched to DeepSeek entirely once I decided to put 10 bucks on it and I realized that it could do whatever I was throwing at Claude or ChatGPT prior to that.

I recommended it to one of my friends, and he was surprised DeepSeek could solve task that Claude got stuck at. I was surprised at it too.

I know others that tried and were less impressed too.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#328

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

> They require constant supervision

I think that may be part of it.

LLMs can be autonomous to an extent. All of them need steering - which is why I feel they are more a superpower the more I am an expert on the subject matter.

The more you want it to be autonomous, than yeah, you may benefit from using the very best the industry has to offer, however slightly better it is.

But if you are always in the loop anyway, you may want to try DeepSeek. You will get similar results for a fraction of the price.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#329

Earlier quoted context omitted.

>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.

Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.

This reply must have cost dozens of dollars.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#330
post #160

Out of curiosity, what's currently the best model I can use locally?

$500k - Kimi K3 (maybe $250k? Haven’t done this one) $25k - DSv4 Flash $4k - Qwen 3.6 35A3B Q5 $1k - Qwen 3.6 27B Q4 Some people prefer the sense over the MoE YMMV.

16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k.

2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k.

1 A4500 can run 35A3B. Those are about $1200 new.

27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.

Post reply on HN