Live data from Hacker News

Qwen3-Coder-Next

qwen.ai

401–410 of 443 posts

Re: Qwen3-Coder-Next

#401

Earlier quoted context omitted.

> I experimented with the Q2 and Q4 quants. Of course you get degraded performance with this.

Obviously. That's why I led with that statement. Those are the quant thresholds where people with mid-high end hardware can run this locally at reasonable speed, though. In my experience Q2 is flakey, but Q4 isn't dramatically worse.

> Obviously. That's why I led with that statement.

Then why did you write this?

> It's always possible that there are some bugs in early implementations that need to be fixed later, but so far I don't see any reason to believe this is actually a Sonnet 4.5 level model.

Re: Qwen3-Coder-Next

#402

Earlier quoted context omitted.

> I do not want my career to become dependent upon Anthropic As someone who switches between Anthropic and ChatGPT depending on the month and has dabbled with other providers and some local LLMs, I think this fear is unfounded. It's really easy to switch between models. The different models have some differences that you notice over time but the techniques you learn in one place aren't going to lock you into a provid…

> It's really easy to switch between models. The different models have some differences that you notice over time but the techniques you learn in one place aren't going to lock you into a provider anywhere. We have two cell phone providers. Google is removing the ability to install binaries, and the other one has never allowed freedom. All computing is taxed, defaults are set to the incumbent monopolies. Searching, e…

I just don’t see it.

I mean, the long arch of computing history has had us wobble back and forth in regards to how closed down it all was, but it seems we are almost at a golden age again with respect to good enough (if not popular) hardware.

On the software front, we definitely swung back from the age of Microsoft. Sure, Linux is a lot more corporate than people admit, but it’s a lot more open than Microsoft’s offerings and it’s capable of running on practically everything except the smallest IOT device.

As for LLMs. I know people have hyped themselves up to think that if you aren’t chasing the latest LLM release and running swarms of agents, you are next in the queues for the soup kitchens, but again, I don’t see why it HAS to play out that way, partly because of history (as referenced), partly because open models are already so impressive and I don’t see any reason why they wouldn’t continue to do well.

In fact, I do my day-to-day work using an open weight model. Beyond that, can only say I know employers who will probably never countenance using commercially hosted LLMs, but who are already setting up self-hosted ones based on open weight releases.

Re: Qwen3-Coder-Next

#404

This is model 12188, which claims to rival SOTA models while not even being in the same league. In terms of intelligence per compute, it’s probably the best model I can realistically run locally on my laptop for coding. It’s solid for scripting and small projects. I tried it on a mid-size codebase (~50k LOC), and the context window filled up almost immediately, making it basically unusable unless you’re extremely exp…

What are you talking about? Qwen3-Coder-Next supports 256k context. Did you wanted to say that you don't have enough memory to run it locally yourself?

Re: Qwen3-Coder-Next

#405
post #3

This GGUF is 48.4GB - https://huggingface.co/Qwen/Qwen3-Coder-Next-GGUF/tree/main/... - which should be usable on higher end laptops. I still haven't experienced a local model that fits on my 64GB MacBook Pro and can run a coding agent like Codex CLI or Claude code well enough to be useful. Maybe this will be the one? This Unsloth guide from a sibling comment suggests it might be: https://unsloth.ai/docs/models/qwen3…

We need a new word, not "local model" but "my own computers model" CapEx based This distinction is important because some "we support local model" tools have things like ollama orchestration or use the llama.cpp libraries to connect to models on the same physical machine. That's not my definition of local. Mine is "local network". so call it the "LAN model" until we come up with something better. "Self-host" exists b…

Local as in localhost

Re: Qwen3-Coder-Next

#406
post #259

Earlier quoted context omitted.

I experimented with the Q2 and Q4 quants. First impression is that it's amazing we can run this locally, but it's definitely not at Sonnet 4.5 level at all. Even for my usual toy coding problems it would get simple things wrong and require some poking to get to it. A few times it got stuck in thinking loops and I had to cancel prompts. This was using the recommended settings from the unsloth repository. It's always p…

I would not go below q8 if comparing to sonnet.

Yeah. Q2 in any model is just severely damaged, unfortunately. Wish it weren’t so.

Re: Qwen3-Coder-Next

#407

Earlier quoted context omitted.

> It's really easy to switch between models. The different models have some differences that you notice over time but the techniques you learn in one place aren't going to lock you into a provider anywhere. We have two cell phone providers. Google is removing the ability to install binaries, and the other one has never allowed freedom. All computing is taxed, defaults are set to the incumbent monopolies. Searching, e…

I just don’t see it. I mean, the long arch of computing history has had us wobble back and forth in regards to how closed down it all was, but it seems we are almost at a golden age again with respect to good enough (if not popular) hardware. On the software front, we definitely swung back from the age of Microsoft. Sure, Linux is a lot more corporate than people admit, but it’s a lot more open than Microsoft’s offer…

> but it seems we are almost at a golden age again with respect to good enough (if not popular) hardware.

I don't think we're in any golden age since the GPU shortages started, and now memory and disks are becoming super expensive too.

Hardware vendors have shown they don't have an interest in serving consumers and will sell out to hyperscalers the moment they show some green bills. I fear a day where you won't be able to purchase powerful (enough) machines and will be forced to subscribe to a commercial provider to get some compute to do your job.

Re: Qwen3-Coder-Next

#409
post #12
post #3

This GGUF is 48.4GB - https://huggingface.co/Qwen/Qwen3-Coder-Next-GGUF/tree/main/... - which should be usable on higher end laptops. I still haven't experienced a local model that fits on my 64GB MacBook Pro and can run a coding agent like Codex CLI or Claude code well enough to be useful. Maybe this will be the one? This Unsloth guide from a sibling comment suggests it might be: https://unsloth.ai/docs/models/qwen3…

I run Qwen3-Coder-30B-A3B-Instruct gguf on a VM with 13gb RAM and a 6gb RTX 2060 mobile GPU passed through to it with ik_llama, and I would describe it as usable, at least. It's running on an old (5 years, maybe more) Razer Blade laptop that has a broken display and 16gb RAM. I use opencode and have done a few toy projects and little changes in small repositories and can get pretty speedy and stable experience up to…

30-A3B model gives 13 t/s without GPU (I noticed that token/sec * # of params matches memory bandwidth).

Re: Qwen3-Coder-Next

#410
post #328

Earlier quoted context omitted.

Who cares? If you don't like it, you can fine tune.

I think a lot of people care. Most decidedely not you.

I think people care about open weights, so they can use it locally, including fine tuning like unalignment.

There are of course people that when you give them something that did cost millions of dollars to build for free will complain and share with the world what exactly they're entitled to.

Post reply on HN