Live data from Hacker News

Qwen 3.8

twitter.com

371–380 of 793 posts

Re: Qwen 3.8

#371

Earlier quoted context omitted.

Any open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC. The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful. Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per se…

I can only imagine what that does to the poor SSD

Paging in experts is mostly reading so on the first order effects it would be fine. These might be second order effects from e.g. write caches needing to be flushed more often (and maybe swapping other applications, if you use a swap file or partition) but it probably wouldn't be too much of an issue.

Re: Qwen 3.8

#372
post #19

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

This seems more of a battle for frontier AI supremacy. I'm afraid that small capable models have been left in the dust. Big labs don't really want to hand over the golden eggs goose to the end user. Possibly the hardware vendors(e.g. Nvidia) may want to play in that area as well, to pull money from all parties.

> I'm afraid that small capable models have been left in the dust

I wholly disagree. Rather than going the "everything is a claude code skill" route, I've been hacking together purpose-built harnesses for all sorts of tasks, and in that environment a wee little baby model can do some really useful things. You end up burning lots of tokens making the thing, but then all that investment comes back when the resulting tool works perfectly fine on a dinky little model that fits on my 3060 Ti.

Re: Qwen 3.8

#373
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

> It's hard to say what their motivation is. Not for anyone who reads history. Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because th…

Sums up exactly what is going on with China's hundred years strategy with their pure focus on technology.

History doesn't repeat itself but it does rhyme.

Re: Qwen 3.8

#374
post #125

Earlier quoted context omitted.

> Anthropic should not have bugged their knowledge distillation attacks. It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)

Appealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.

[flagged]

Re: Qwen 3.8

#375
post #4

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…

These things take months to train. No chance this is a reaction to what just happened.

Re: Qwen 3.8

#376
post #6
post #4

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

> It's hard to say what their motivation is.

Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].

The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.

I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.

1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...

Re: Qwen 3.8

#378

Earlier quoted context omitted.

My experience was so much different to this, that I have the unfortunate impression that you're shilling. It really was not a capable model, it felt like the old oai models back when we were all excited but couldn't actually trust them even in the littlest ways. What harness were you using, did you do any work to make it better? What was I doing wrong? I just pointed opencode at it, with a pretty simple (large-ish) d…

It’s so fascinating watching people on here bicker about models like fine wines. Wild times we live in. Wild times.

It's not a frontier LLM if it's not made in the Silicon Valley region of California, otherwise it's just a sparkling LLM.

Re: Qwen 3.8

#379
post #324
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

> It's hard to say what their motivation is. Feels pretty easy to me. They want to turn LLMs into a commodity, and watch the US AI labs crash and burn. There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens an…

Or it could be good for humanity.

Re: Qwen 3.8

#380

Earlier quoted context omitted.

Please don't conflate a volunteer effort with no expected economic gain with a very well funded company (or fleet of companies). With the CCP's highly successful track record with subsuming other markets, Occam's razor applies to why they're doing this.

[flagged]

Well of course nobody believes that, there are a large number of great Chinese open source projects, and plenty of great Chinese contributors to open source projects. I have zero doubts that Chinese people are at least equally as capable of embracing open source as anyone else. It would be strange to suggest otherwise.

That said, though, I do have trouble believing the long-term story for open weights, anywhere. We do not need an evil government for open weight to "make sense", but I do think we need some government involvement for open weight to make sense in the long run. Otherwise, it's not 100% clear how they could be sustainable, and I don't think massive companies really can be trusted to just be philanthropic with no incentives indefinitely (or really, much at all to begin with.)

Chinese models being open weight does help them gain some Western mindshare, whereas for obvious reasons Americans would be very suspicious of running their source code and prompts through Chinese providers. (And I think that's justified, I just also think that American providers aren't really that much better in the long run, and you should prefer to not have to go through any provider for true privacy.)

Post reply on HN