Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
qwenlm.github.io
Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
1–10 of 32 posts
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#2A Chinese company announcing this on Spring Festival eve, that is very surprising. The deep seek announcement must have put a fire under them. I am surprised anything is being done right now in these Chinese tech companies.
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#3Party goes on
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#4[deleted]
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#5Now they need to finetune it like R1 and o1 and it will be competitive with SOTA models.
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#6This appears to be Qwen's new best model, API only for the moment, which they say is better than DeepSeek v3.
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#7A Chinese company announcing this on Spring Festival eve, that is very surprising. The deep seek announcement must have put a fire under them. I am surprised anything is being done right now in these Chinese tech companies.
It's like when Gemini topped Chatbot Arena Leaderboard, and OpenAI released a model next day.
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#8This appears to be Qwen's new best model, API only for the moment, which they say is better than DeepSeek v3.
It is available at https://chat.qwenlm.ai/, under the model selector.
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#9A Chinese company announcing this on Spring Festival eve, that is very surprising. The deep seek announcement must have put a fire under them. I am surprised anything is being done right now in these Chinese tech companies.
Well, DeepSeek engineers are (desperately) fire-fighting as they don't have nearly as much capacity as needed. Competitors either already rushed release or decided to do an hush release of whatever they had in the pipeline. Sounds like everyone is working L
Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model
#10HuggingFace demo: https://huggingface.co/spaces/Qwen/Qwen2.5-Max-Demo
Source: https://x.com/Alibaba_Qwen/status/1884263157574820053