Live data from Hacker News

Qwen 3.8

twitter.com

41–50 of 793 posts

Re: Qwen 3.8

#41

The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.

5.6-sol would be a better comparison given it's general availability and usage allowances

And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.

Re: Qwen 3.8

#42

Just imagine Anthropic making Opus open-weights now for the sake of trolling everyone. Wouldn't surprise me at this point xD

That OpenAI releases Sol as downloadable weights feels way more likely than Anthropic releasing even the tiniest of models for download.

[deleted]

Re: Qwen 3.8

#43

Qwen is the most censored of the Chinese models in my testing, which makes me wonder in what other ways it is compromised. Open weights doesn't really reveal what's in there. And, in my tests, existing Qwen models are not at the pareto frontier of any metric; DeepSeek V4 Pro is better, faster, and much cheaper than Qwen 3.7 Max. (DeepSeek is also among the least censored of the Chinese models.) I guess we'll see if t…

What censorship? ;) https://github.com/p-e-w/heretic

Re: Qwen 3.8

#44
post #26
post #6

Earlier quoted context omitted.

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along. So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue. As an example, HuggingFace recently was investigating a security incid…

How does this explain open weights? They could easily take the same closed route like their American friends

Re: Qwen 3.8

#46

Earlier quoted context omitted.

Humanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.

I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.

The fear is not about the models open weights it is the erosion of training capability in other countries. Why train models when they do it for free? Until they don't of course, or they start doing what the US is doing right now by locking out some models to government only or internal market only.

What people should be afraid is the rug pull.

Re: Qwen 3.8

#47
I do hope they provide optimized A3B quants--that's been the sweet spot for usable local inference for me, at least.

Re: Qwen 3.8

#48
post #4

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…

It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.

Re: Qwen 3.8

#49
Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.

Re: Qwen 3.8

#50
post #32
post #19

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

I think everyone is hoping this! It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.

Indeed! That would be the sweet spot for my 2x3090 rig
Post reply on HN