Live data from Hacker News

Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

qwenlm.github.io

11–20 of 32 posts

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#11
post #9
post #2

A Chinese company announcing this on Spring Festival eve, that is very surprising. The deep seek announcement must have put a fire under them. I am surprised anything is being done right now in these Chinese tech companies.

Well, DeepSeek engineers are (desperately) fire-fighting as they don't have nearly as much capacity as needed. Competitors either already rushed release or decided to do an hush release of whatever they had in the pipeline. Sounds like everyone is working L

They are being attacked as well.

https://apnews.com/article/deepseek-ai-artificial-intelligen...

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#12
post #7
post #2

A Chinese company announcing this on Spring Festival eve, that is very surprising. The deep seek announcement must have put a fire under them. I am surprised anything is being done right now in these Chinese tech companies.

It's like when Gemini topped Chatbot Arena Leaderboard, and OpenAI released a model next day.

is gemini really better than e.g. claude 3.5?

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#16
> We evaluate Qwen2.5-Max alongside leading models

> [...] we are unable to access the proprietary models such as GPT-4o and Claude-3.5-Sonnet. Therefore, we evaluate Qwen2.5-Max against DeepSeek V3

"We'll compare our proprietary model to other proprietary models. Except when we don't. Then we'll compare to non-proprietary models."

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#17
post #12
post #7

Earlier quoted context omitted.

It's like when Gemini topped Chatbot Arena Leaderboard, and OpenAI released a model next day.

is gemini really better than e.g. claude 3.5?

Mostly, but not for coding. Also the arena is pretty much a vibes benchmark. Don't take it too seriously. Livebench is a better indicator

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#18

The significance of _all_ of these releases at once is not lost on me. But the reason for it is lost on me. Is there some convention? Is this political? Business strategy?

Today is the last day before the Chinese New Year.

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#19
>Many critical details regarding this scaling process were only disclosed with the recent release of DeepSeek V3

And so they decide to not disclose their own training information just after they told everyone how useful it was to get Deepseeks? Honestly can't say I care about "nearly as good as o1" when its a closed API with no additional info.

Post reply on HN