Earlier quoted context omitted.
Any open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC. The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful. Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per se…
I can only imagine what that does to the poor SSD
Qwen 3.8
371–380 of 793 posts
Re: Qwen 3.8
#372Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data
This seems more of a battle for frontier AI supremacy. I'm afraid that small capable models have been left in the dust. Big labs don't really want to hand over the golden eggs goose to the end user. Possibly the hardware vendors(e.g. Nvidia) may want to play in that area as well, to pull money from all parties.
I wholly disagree. Rather than going the "everything is a claude code skill" route, I've been hacking together purpose-built harnesses for all sorts of tasks, and in that environment a wee little baby model can do some really useful things. You end up burning lots of tokens making the thing, but then all that investment comes back when the resulting tool works perfectly fine on a dinky little model that fits on my 3060 Ti.
Re: Qwen 3.8
#373Earlier quoted context omitted.
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
> It's hard to say what their motivation is. Not for anyone who reads history. Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because th…
History doesn't repeat itself but it does rhyme.
Re: Qwen 3.8
#374Earlier quoted context omitted.
> Anthropic should not have bugged their knowledge distillation attacks. It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)
Appealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.
Re: Qwen 3.8
#375I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
Re: Qwen 3.8
#376I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].
The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.
I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.
1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...
Re: Qwen 3.8
#377Re: Qwen 3.8
#378Earlier quoted context omitted.
My experience was so much different to this, that I have the unfortunate impression that you're shilling. It really was not a capable model, it felt like the old oai models back when we were all excited but couldn't actually trust them even in the littlest ways. What harness were you using, did you do any work to make it better? What was I doing wrong? I just pointed opencode at it, with a pretty simple (large-ish) d…
It’s so fascinating watching people on here bicker about models like fine wines. Wild times we live in. Wild times.
Re: Qwen 3.8
#379Earlier quoted context omitted.
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…
Re: Qwen 3.8
#380Earlier quoted context omitted.
Please don't conflate a volunteer effort with no expected economic gain with a very well funded company (or fleet of companies). With the CCP's highly successful track record with subsuming other markets, Occam's razor applies to why they're doing this.
[flagged]
That said, though, I do have trouble believing the long-term story for open weights, anywhere. We do not need an evil government for open weight to "make sense", but I do think we need some government involvement for open weight to make sense in the long run. Otherwise, it's not 100% clear how they could be sustainable, and I don't think massive companies really can be trusted to just be philanthropic with no incentives indefinitely (or really, much at all to begin with.)
Chinese models being open weight does help them gain some Western mindshare, whereas for obvious reasons Americans would be very suspicious of running their source code and prompts through Chinese providers. (And I think that's justified, I just also think that American providers aren't really that much better in the long run, and you should prefer to not have to go through any provider for true privacy.)