Live data from Hacker News

Qwen3.6-Plus: Towards real world agents

qwen.ai

91–100 of 235 posts

Re: Qwen3.6-Plus: Towards real world agents

#93
post #59
post #17

Earlier quoted context omitted.

> I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. There isn't, pretty much everyone wants the best of the best.

No. Right now I'm upset that Google has removed (or at least is in the process of removing) the Gemini 2.0 flash model. We use it for some pretty basic functionality because it's cheap and fast and honestly good enough for what we use it for in that part of our app. We're being forced to "upgrade" to models that are at least 2.5 times as expensive, are slower and, while I'm sure they're better for complex tasks, don'…

this is one of the reasons im hearing more and more people are using open/locally hosted models. particularly so we dont have to waste time to entirely redo everything when inevitably a company decides to pull the rug out from under us and change or remove something integral to our flow, which over the years we've seen countless times, and seems to be getting more and more common.

products entirely disappearing or significantly changing will be more and more common in the llm arena as things move forward towards companies shutting down, bubbles deflating, brand priorities drastically reshifting, etc...

i think, we're at or at least close to a time to really put some thought into which pieces of your flow could be done entirely with an open/local model and be honest with ourselves on which pieces of our flow truly needs sota or closed models that may entirely disappear or change. in the long run, putting a little bit of thought into this now will save a lot of headache later.

Re: Qwen3.6-Plus: Towards real world agents

#94

This is their hosted-only model, not an open weight model like they’ve become known for. They got a lot of good publicity for their open weight model releases, which was the goal. The hard part is pivoting from an open weight provider to being considered as a competitor to Claude and ChatGPT. Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as…

I’m starting to wonder where the most is for any of these models.

Sure they are not cheap to train. But if open weight models continue to be trained and continue to become available on cheaper hardware, how do dedicated AI companies protect their margins?

Re: Qwen3.6-Plus: Towards real world agents

#95

I understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or…

Opus 4.5 is already pretty good.

Opus 4.5 is $25/m output tokens.

This is at most $6/m output tokens.

That's ~1/4 the price.

Re: Qwen3.6-Plus: Towards real world agents

#96

This is their hosted-only model, not an open weight model like they’ve become known for. They got a lot of good publicity for their open weight model releases, which was the goal. The hard part is pivoting from an open weight provider to being considered as a competitor to Claude and ChatGPT. Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as…

> not an open weight model like they’ve become known for. Right, they state that they'll release "smaller" variants openly at some point, with few details as to what that means. Will there be a ~300B variant as with Qwen 3.5? The blog post doesn't say.

I'm not interested in adopting an inferior closed source weight from a geopolitical rival. The open source weights argument was the one thing China had going and that I was seriously cheering them on for. They could have been our saviors and disrupted the US tech giants - and if it was open, I'd have welcomed it.

Now they show their true colors. They want to train models on our engineering to replace us, while simultaneously giving nothing back? No thanks. I'd rather fund the shitty US hyperscalers. At least that leads to jobs here.

If there's a company willing develop and foster large scale weights in the open, I'll adopt their tooling 100%. It doesn't matter if they're a year behind. Just do it open and build an entire ecosystem on top of it.

The re-AOLization of the internet into thin clients is bullshit, and all it takes is one player to buck the rules to topple the whole house of cards.

Re: Qwen3.6-Plus: Towards real world agents

#97
post #93
post #59

Earlier quoted context omitted.

No. Right now I'm upset that Google has removed (or at least is in the process of removing) the Gemini 2.0 flash model. We use it for some pretty basic functionality because it's cheap and fast and honestly good enough for what we use it for in that part of our app. We're being forced to "upgrade" to models that are at least 2.5 times as expensive, are slower and, while I'm sure they're better for complex tasks, don'…

this is one of the reasons im hearing more and more people are using open/locally hosted models. particularly so we dont have to waste time to entirely redo everything when inevitably a company decides to pull the rug out from under us and change or remove something integral to our flow, which over the years we've seen countless times, and seems to be getting more and more common. products entirely disappearing or si…

What’s interesting about this is that for previous technologies you could define a standard and demonstrate compliance with interfaces and behavior.

But with LLMs, how do you know switching from one to another won’t change some behavior your system was implicitly relying on?

Re: Qwen3.6-Plus: Towards real world agents

#98
post #17

Earlier quoted context omitted.

> I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. There isn't, pretty much everyone wants the best of the best.

> There isn't, pretty much everyone wants the best of the best. For direct user interaction or coding problems, perhaps. But as API calls get cheaper, it becomes more realistic to use them for completely automated workflows against data-sets, or as sub-agents called from expensive SOTA models. For example, in Claude, using Opus as an orchestrator to call Sonnet sub-agents, is a popular usage "hack." That only gets mo…

That is a very complex, high level use case that takes time to configure and orchestrate.

There are many simpler tasks that would work fine with a simpler, local model.

Re: Qwen3.6-Plus: Towards real world agents

#99
post #63

Earlier quoted context omitted.

I think it’s more the principle of deception that upsets people. Imagine if Apple released a new iPhone and publicly compared its specs to some previous gen Android. It’s not in good faith.

They compared their M-series chips to older Intel Macs for a while, likely to target users who were still on Intel chips. If they released a lower cost iPhone and compared it to a previous gen Android I could see the reasoning for it. It's not deception if it's a valid comparison and people just fail to understand what's being compared. Now, is it mildly deceptive because all of the companies using incredibly confusi…

Apple continues to compare to prior versions of Apple Silicon. I suspect it is a mix of trying to provide useful, realistic upgrade information and numbers that still sound good for those not paying attention.

I don't think any org doing this is necessarily being deceptive, so long as there's some reasonable basis for the chosen comparable(s).

For example, comparing a new iPhone to a prior Android phone might make sense if the install base is considerably large and Apple is targeting the cohort for user acquisition. (~"These benchmarks are not for you.")

The community will always run the numbers and get the clicks for the benchmarks not filled in by the 1st party. I noticed what appeared to be some movement from Apple in content they've produced to get ahead of this with recent product content.

Re: Qwen3.6-Plus: Towards real world agents

#100
post #34

Earlier quoted context omitted.

The OpenRouter usage stats indicate the opposite: https://openrouter.ai/rankings?view=month

OpenRouter usage is likely skewed towards LLMs that are more niche and/or self-hostable by solid hardware that's available, but most consumers don't have on hand. I can imagine Anthropic and OpenAI LLMs often get called directly from their APIs instead. At least from my experience and friends of mine, we use OpenRouter for cases where we want to use smaller LLMs like Qwen, but when I've used ChatGPT and Claude, I use…

I use ChatGPT and Claude on OpenRouter, because it's just easier than buying credits on each platform separately.
Post reply on HN