Live data from Hacker News

Qwen3.6-Plus: Towards real world agents

qwen.ai

61–70 of 235 posts

Re: Qwen3.6-Plus: Towards real world agents

#61
post #2

The benchmarks provided are for Opus-4.5, not for the latest Opus-4.6 and Qwen is still lagging in a lot of them.

And it seems they've decided to go closed-source for their largest, best models.

They always did that. Did they say anywhere they'd open all their models? They still have a business.

Re: Qwen3.6-Plus: Towards real world agents

#62
post #5
post #2

The benchmarks provided are for Opus-4.5, not for the latest Opus-4.6 and Qwen is still lagging in a lot of them.

There is no reason to benchmark against Opus 4.5 when Opus 4.6 has been out so long, other than to be misleading.

I can see reasons, among others that 4.5 was the one established as they were preparing this version. "So long" is merely 2 months ago, and Qwen 3.5 was barely released less than 2 months ago. They were likely already working on finalizing 3.6 before 3.5 official launch, and as 4.6 came out.

In any case, aside Claude fanboyism, having other plays inch closer to similar performance is always useful. Even if they are "6 months behind" as the pace slows down, this guarantees that there's no huge moat and they'll eventually either get to where the SOTA is, or the difference wont be that big.

I'd rather put fewer eggs in 2-3 big player baskets.

Re: Qwen3.6-Plus: Towards real world agents

#63

I understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or…

I think it’s more the principle of deception that upsets people. Imagine if Apple released a new iPhone and publicly compared its specs to some previous gen Android. It’s not in good faith.

Re: Qwen3.6-Plus: Towards real world agents

#64

This is their hosted-only model, not an open weight model like they’ve become known for. They got a lot of good publicity for their open weight model releases, which was the goal. The hard part is pivoting from an open weight provider to being considered as a competitor to Claude and ChatGPT. Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as…

> Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as advertising, not because they were feeling generous.

The naivety around this has been staggering quite frankly. All of a sudden, people thinking that meta etc are releasing free models because they believe in open access and distribution of knowledge. No, they just suck comparatively. There is nothing to sell. Using it to recruit and generate attention is the best play for them.

Re: Qwen3.6-Plus: Towards real world agents

#65
post #3

Worth noting that this model, unlike almost all qwen models, is not open-weight, nor is the parameter count exposed. Also odd that it is compared against opus 4.5 even though 4.6 was released like 2 months ago.

If Opus 4.6 was only released two months ago, then it seems reasonable that Qwen hasn't finished fully comparing against the latest Opus.

Re: Qwen3.6-Plus: Towards real world agents

#66
Pretty solid Pelican: https://gist.github.com/simonw/ca081b679734bc0e5997a43d29fad...

I used the https://modelstudio.alibabacloud.com/ API to generate that one, which required signing up for an account and attaching PayPal billing - but it looks like OpenRouter are offering it for free right now so I could have used that: https://openrouter.ai/qwen/qwen3.6-plus:free

Re: Qwen3.6-Plus: Towards real world agents

#69
post #17

Earlier quoted context omitted.

> I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. There isn't, pretty much everyone wants the best of the best.

The OpenRouter usage stats indicate the opposite: https://openrouter.ai/rankings?view=month

what happened around jan this year(26) that caused such a climb in usage?

Re: Qwen3.6-Plus: Towards real world agents

#70
post #63

I understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or…

I think it’s more the principle of deception that upsets people. Imagine if Apple released a new iPhone and publicly compared its specs to some previous gen Android. It’s not in good faith.

Why are we so quick to call it deception? Their figure is quite clear. They aren't fiddling with the graph or hiding the labels, they are clearly stating which models it compares against. But I agree on the sentiment that the standard practice should be to bench against the latest SOTA models.
Post reply on HN