Go China, screw America* *within the scope of open models only
I like my Apache 2.0 licensed Gemma, and NVIDIA’s Nemotrons are decent bases for finetuning or continued pretraining, esp thanks to good documentation and tooling. Oh, and Mira’s thinking machines lab dropped Inkling, a ~1T open weight model too. This isn’t US vs China. This is open vs closed.
Qwen 3.8
71–80 of 793 posts
Re: Qwen 3.8
#72Just imagine Anthropic making Opus open-weights now for the sake of trolling everyone. Wouldn't surprise me at this point xD
That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.
Re: Qwen 3.8
#73I have a feeling this is the next…frontier of that fight
One can only hope it eventually does as well as Linux
Re: Qwen 3.8
#74I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
Re: Qwen 3.8
#75Re: Qwen 3.8
#76I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.
Re: Qwen 3.8
#77Earlier quoted context omitted.
I would rather see them releasing 3.7-27B, 3.7-122B or their 3.8 versions. Qwen/QwQ were always about the best available local inference at home.
I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.
You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.
Re: Qwen 3.8
#78I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.
IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?
Re: Qwen 3.8
#79Earlier quoted context omitted.
I would rather see them releasing 3.7-27B, 3.7-122B or their 3.8 versions. Qwen/QwQ were always about the best available local inference at home.
I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.
Re: Qwen 3.8
#80The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.
I wouldn't call it a moat, but I would call it a noticeably better model. Subjectively, for my own work, I would rate the top models Fable > K3 > Sol. But it's not like Fable is so substantially better than the other two that I would be seriously impacted if I didn't have access to it anymore. All three are amazing models, and of the three, Fable is the only one that regularly triggers refusals.
Coordinating agents though? Fable any day.