The benchmarks provided are for Opus-4.5, not for the latest Opus-4.6 and Qwen is still lagging in a lot of them.
And it seems they've decided to go closed-source for their largest, best models.
Qwen3.6-Plus: Towards real world agents
61–70 of 235 posts
Re: Qwen3.6-Plus: Towards real world agents
#62The benchmarks provided are for Opus-4.5, not for the latest Opus-4.6 and Qwen is still lagging in a lot of them.
There is no reason to benchmark against Opus 4.5 when Opus 4.6 has been out so long, other than to be misleading.
In any case, aside Claude fanboyism, having other plays inch closer to similar performance is always useful. Even if they are "6 months behind" as the pace slows down, this guarantees that there's no huge moat and they'll eventually either get to where the SOTA is, or the difference wont be that big.
I'd rather put fewer eggs in 2-3 big player baskets.
Re: Qwen3.6-Plus: Towards real world agents
#63I understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or…
Re: Qwen3.6-Plus: Towards real world agents
#64This is their hosted-only model, not an open weight model like they’ve become known for. They got a lot of good publicity for their open weight model releases, which was the goal. The hard part is pivoting from an open weight provider to being considered as a competitor to Claude and ChatGPT. Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as…
The naivety around this has been staggering quite frankly. All of a sudden, people thinking that meta etc are releasing free models because they believe in open access and distribution of knowledge. No, they just suck comparatively. There is nothing to sell. Using it to recruit and generate attention is the best play for them.
Re: Qwen3.6-Plus: Towards real world agents
#65Worth noting that this model, unlike almost all qwen models, is not open-weight, nor is the parameter count exposed. Also odd that it is compared against opus 4.5 even though 4.6 was released like 2 months ago.
Re: Qwen3.6-Plus: Towards real world agents
#66I used the https://modelstudio.alibabacloud.com/ API to generate that one, which required signing up for an account and attaching PayPal billing - but it looks like OpenRouter are offering it for free right now so I could have used that: https://openrouter.ai/qwen/qwen3.6-plus:free
Re: Qwen3.6-Plus: Towards real world agents
#67Re: Qwen3.6-Plus: Towards real world agents
#68Re: Qwen3.6-Plus: Towards real world agents
#69Earlier quoted context omitted.
> I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. There isn't, pretty much everyone wants the best of the best.
The OpenRouter usage stats indicate the opposite: https://openrouter.ai/rankings?view=month
Re: Qwen3.6-Plus: Towards real world agents
#70I understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or…
I think it’s more the principle of deception that upsets people. Imagine if Apple released a new iPhone and publicly compared its specs to some previous gen Android. It’s not in good faith.