Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

221–230 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#221
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

the ceiling is to eval's quality

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#222

Earlier quoted context omitted.

It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.

VCs are footing the bill for that $200 subscription.

The $200 sub is customer acquisition cost to hook devs that then become the marketing team trying to get their company to bring in Claude (at the highly profitable API price).

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#223

Earlier quoted context omitted.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

If the legal system declares the first thief’s theft not theft then all bets are off.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#224

Earlier quoted context omitted.

But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.

Effective jargon usage is understood by the target audience. If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.

Or you're below its level.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#225

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.

yeah, i tried out GLM-5.2 when the news was all full of hype for that, and it's fine... definitely better value that API rates for claude. but comparing the value i got from that to the value i get from a claude max subscription... claude is way cheaper.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#226

Earlier quoted context omitted.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

If the legal system declares the first thief’s theft not theft then all bets are off.

> If the legal system declares the first thief’s theft not theft

But they didn't find it. The Big LLM provider accepted guilt and paid a fine.

You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#227
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

Don't say that too loud, you may burst the bubble prematurely.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#228

Earlier quoted context omitted.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

If the legal system declares the first thief’s theft not theft then all bets are off.

Is it theft if another thief steal's the first thief's theft?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#229
post #5

I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth. I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A vers…

Did you also try Opus 5 and 5.6 Sol?

5.6 sol was very impressive for me. I had a weird behavior while using Qt and I gave it a screenshot and my expectation of what should happen and it read the Qt sourcode and showed me that my issue was a bug (including link to the ticket).

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#230
post #207

Earlier quoted context omitted.

Is there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?

Opencode Go has Qwen 3.8 Max at $10/month

Found the usage limits on OpenCode Go quite poor tbh
Post reply on HN