Live data from Hacker News

QwQ-32B: Embracing the Power of Reinforcement Learning

qwenlm.github.io

21–30 of 178 posts

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#21
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no.

I get running these models is not cheap, but they just lost a potential customer / user.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#22

This is ridiculous. 32B and beating deepseek and o1. And yet I'm trying it out and, yeah, it seems pretty intelligent... Remember when models this size could just about maintain a conversation?

I still remember Vicuna-33B, that one stayed on the leaderboards for quite a while. Today it looks like a Model T, with 1B models being more coherent.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#23
post #12

This is insane matching deepseek but 20x smaller?

I wonder if having a big mixture of experts isn't all that valuable for the type of tasks in math and coding benchmarks. Like my intuition is that you need all the extra experts because models store fuzzy knowledge in their feed-forward layers, and having a lot of feed-forward weights lets you store a longer tail of knowledge. Math and coding benchmarks do sometimes require highly specialized knowledge, but if we bel…

MoE is likely temporary, local optimum now that resembles bitter lesson path. With the time we'll likely distill what's important, shrink it and keep it always active. There may be some dynamic retrieval of knowledge (but not intelligence) in the future but it probably won't be anything close to MoE.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#27
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.

[deleted]

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#28
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.

You can also try the HuggingFace Space at https://huggingface.co/spaces/Qwen/QwQ-32B-Demo (though it seems to be fully utilized at the moment)

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#29
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

super impressive. we won't need that many GPUs in the future if we can have the performance of DeepSeek R1 with even less parameters. NVIDIA is in trouble. We are moving towards a world of very cheap compute: https://medium.com/thoughts-on-machine-learning/a-future-of-...

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#30
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.

Running this model is dirt cheap, they're just not chasing that type of customer.
Post reply on HN