Earlier quoted context omitted.
[flagged]
Sir, this is a multi-billion dollar operation. There have to be some incentives.
Qwen 3.8
441–450 of 793 posts
Re: Qwen 3.8
#442Earlier quoted context omitted.
Who do you buy DeepSeek from? I bought it through OpenRouter and used it with Pi agent. The model was good, but there appeared to be a pricing glitch or something, because it burned through $50 in under an hour on pretty trivial stuff. Pi agent claimed it only used like $1. OpenRouter claimed differently and said I used all $50.
I can highly recommend OpenCode Go. I use it from pi.dev as well through the OpenCode Go $10 subscription ($5 first month). Used more than 20M tokens at a cost of ~$20 (up to $60 is included in the $5 plan) Out of which deepseek pro had ~200 messages which is around 1.5M tokens (10+M cached)
Re: Qwen 3.8
#443I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
So what would the long game be for chinese companies?
Re: Qwen 3.8
#444It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
Re: Qwen 3.8
#445> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
Re: Qwen 3.8
#446I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
https://en.wikipedia.org/wiki/World_Artificial_Intelligence_...
Re: Qwen 3.8
#447> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
Re: Qwen 3.8
#448I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July. Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8. I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to bett…
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies
Re: Qwen 3.8
#449> "compatible" instead of "comparable." It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.
Re: Qwen 3.8
#450I counted the other day and there were at least 12 different providers with "better than Opus 4.5 performance" on Artificial Analysis, Opus 4.5 being Anthropic's December release that many say kicked off the latest acceleration. Which is totally insane competition, particularly given how low switching costs. I personally think that Opus 4.5 level performance is sufficient for most apps and usecases, as they get deplo…
> Which is totally insane competition, particularly given how low switching costs Which is why OAI and Anthropic will most probably push for more governmental control and bans. Without it their whole income model is cooked.
Then again, all of Chinese models are open. And DeepSeek even publishes research papers alongside their models that go in depth into the methodology. I guess there's not much stopping USian companies from copying