Live data from Hacker News

Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

qwenlm.github.io

21–30 of 32 posts

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#21
I thought there were three DeepSeek items on the HN front page, but this turned out to be a fourth one, because it's the Qwen team saying they have a secret version of Qwen that's actually better than DeepSeek-V3.

I don't remember the last time 20% of the HN front page was about the same thing. Then again, nobody remembers the last time a company's market cap fell by 569 billion dollars like NVIDIA did yesterday.

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#22
post #21

I thought there were three DeepSeek items on the HN front page, but this turned out to be a fourth one, because it's the Qwen team saying they have a secret version of Qwen that's actually better than DeepSeek-V3. I don't remember the last time 20% of the HN front page was about the same thing. Then again, nobody remembers the last time a company's market cap fell by 569 billion dollars like NVIDIA did yesterday.

it's a scaling law for stocks!

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#24
post #12
post #7

Earlier quoted context omitted.

It's like when Gemini topped Chatbot Arena Leaderboard, and OpenAI released a model next day.

is gemini really better than e.g. claude 3.5?

Gemini is actually pretty useless because of its rate limits.

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#25

>Many critical details regarding this scaling process were only disclosed with the recent release of DeepSeek V3 And so they decide to not disclose their own training information just after they told everyone how useful it was to get Deepseeks? Honestly can't say I care about "nearly as good as o1" when its a closed API with no additional info.

It's not even "nearly as good as o1". They only compared to the older 4o.

You can safely assume Qwen2.5-Max will score worse than all of the recent reasoning models (o1, DeepSeek-R1, Gemini 2.0 Flash Thinking).

It'll probably become a very strong model if/when they apply RL training for reasoning. However, all the successful recipes for this are closed source, so it may take some time. They could do SFT based on another model's reasoning chains in the meantime, though the DeepSeek-R1 technical report noted that it's not as good as RL training.

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#26

The significance of _all_ of these releases at once is not lost on me. But the reason for it is lost on me. Is there some convention? Is this political? Business strategy?

Today is the last day before the Chinese New Year.

My thoughts go out to the poor engineers who got put on call because someone scheduled a product release on the day before the biggest holiday of their year.

Re: Qwen2.5-Max: Exploring the intelligence of large-scale MoE model

#30
post #21

I thought there were three DeepSeek items on the HN front page, but this turned out to be a fourth one, because it's the Qwen team saying they have a secret version of Qwen that's actually better than DeepSeek-V3. I don't remember the last time 20% of the HN front page was about the same thing. Then again, nobody remembers the last time a company's market cap fell by 569 billion dollars like NVIDIA did yesterday.

Somehow I failed to notice that 4 ÷ 30 is not 20%. It's more like 13%. That was a dumb mistake.
Post reply on HN