Earlier quoted context omitted.
US AI labs really rub me the wrong way, especially with the doom and scare tactics they use. Both Altman and Dario keep talking about how AI will replace workers and how we should regulate LLMs for national security, Dario’s main point. LLMs are useful. We can all see that in agentic coding. But replacing everyone’s job? Hardly. And what’s with the scare tactic of trying to get the US government to ban foreign models…
Sam drank the "superintelligence" kool aid early on and said 30-40% of jobs could be impacted by AI, but recently admitted he was wrong > “My scorecard, at the highest level, would be we’ve been roughly right on technological predictions and pretty wrong on the social and economic implications” https://www.cxtoday.com/ai-automation-in-cx/sam-altman-softe... I agree re: Dario quietly pushing for government control. He…
Qwen3.8-Max: A New Bar for Coding and Cowork
91–100 of 653 posts
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#92Earlier quoted context omitted.
I use it with a strix halo server. 35B runs stupidly fast. 27B is about 700 TPS prefill and 30 TPS token generation. Which interestedly is about what Kimi K3 gives me depending on provider.
what hardware do you use or recommend for this? never heard of it until today.
I personally have the Framework Desktop, but there's also systems from other brands like Bosgame
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#93Can a model be stripped off anything not relevant to coding and get a lot lighter? Or is that impossible? Just like we have professors with specialisation wondering if AI models can also be so.
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#94i might end up cancelling claude, anybody else thinking of the same ?
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#95I think the window for a ban of open weight models is closing fast so let's hope US administration is going to miss it and we get Fable-level models (at least in some aspects) with open weights without infringing any newly introduced law as a long-term local baseline.
What can the US administration do about it?
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#96I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers... How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better…
Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#97Re: Qwen3.8-Max: A New Bar for Coding and Cowork
#98I'm trying and failing to find value running a potential Qwen 3.8 27b dense model on a 16 core, 128 GB of ram, 2080ti box. Yes, the GPU yells for help, but the problem is that no math works to upgrade this machine even when pouring $200 in rent every month into the large model providers... How are you all justifying economical use of these local models right now? What's the cost efficient way to do this and do better…
You will simply not get more value out of running a local model vs paying for a subscription/API from the cloud in 2026. There is no math that will make local models come out ahead in $/intelligence/token.* The point of local models is privacy, offline use, and maybe no guard rails. * Not talking about enterprises that buy DGX racks and host Chinese models for internal use.