Live data from Hacker News

Qwen3.8-Max: A New Bar for Coding and Cowork

qwen.ai

91–100 of 653 posts

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#91
post #52

Earlier quoted context omitted.

US AI labs really rub me the wrong way, especially with the doom and scare tactics they use. Both Altman and Dario keep talking about how AI will replace workers and how we should regulate LLMs for national security, Dario’s main point. LLMs are useful. We can all see that in agentic coding. But replacing everyone’s job? Hardly. And what’s with the scare tactic of trying to get the US government to ban foreign models…

Sam drank the "superintelligence" kool aid early on and said 30-40% of jobs could be impacted by AI, but recently admitted he was wrong > “My scorecard, at the highest level, would be we’ve been roughly right on technological predictions and pretty wrong on the social and economic implications” https://www.cxtoday.com/ai-automation-in-cx/sam-altman-softe... I agree re: Dario quietly pushing for government control. He…

but this crap may take forever to play out even if the outcome is well-known. Self-driving is "here", it's obvious that once it's cheap enough having a human behind a car wheel or a freight truck wheel is an absurd waste of human life (kinda like digging canals with bare hands instead of an excavator), yet truckers and uber drivers are still employed. But everyone knows the writing is on the wall for them.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#92
post #59

Earlier quoted context omitted.

I use it with a strix halo server. 35B runs stupidly fast. 27B is about 700 TPS prefill and 30 TPS token generation. Which interestedly is about what Kimi K3 gives me depending on provider.

what hardware do you use or recommend for this? never heard of it until today.

Strix Halo is the unified memory platform from AMD. Similar to the DGX Spark from NVIDIA or the M series Macs.

I personally have the Framework Desktop, but there's also systems from other brands like Bosgame

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#93

Can a model be stripped off anything not relevant to coding and get a lot lighter? Or is that impossible? Just like we have professors with specialisation wondering if AI models can also be so.

no, apparently, otherwise we'd already have specialized models. every bit of meaningful human-generated data appears to improve the overall capability of the model.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#95
post #67

I think the window for a ban of open weight models is closing fast so let's hope US administration is going to miss it and we get Fable-level models (at least in some aspects) with open weights without infringing any newly introduced law as a long-term local baseline.

What can the US administration do about it?

What they always do. Send in armed men with guns? Export Controls. Import Controls. National Security Laws.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#96
post #79

I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers... How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better…

If you can do inference on the CPU, drop the GPU : it should be faster.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#98
post #79

I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers... How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better…

You will simply not get more value out of running a local model vs paying for a subscription/API from the cloud in 2026. There is no math that will make local models come out ahead in $/intelligence/token.* The point of local models is privacy, offline use, and maybe no guard rails. * Not talking about enterprises that buy DGX racks and host Chinese models for internal use.

Points are starting to be made in favor of value, to the contrary of what you are affirming. Specifically because the new open weights models lower the TCO of hardware in an environment where new open weights were previously thought to be a thing of the past.

Re: Qwen3.8-Max: A New Bar for Coding and Cowork

#99
post #48
post #7

Earlier quoted context omitted.

That's the thing. Wny are companies like OpenAI/Anthropic/Alibaba/Kimi/Deepseek still hiring SWEs if their models have become so good?

https://en.wikipedia.org/wiki/Jevons_paradox

Otherwise known as the "no shit Sherlock" principle
Post reply on HN