To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
I get running these models is not cheap, but they just lost a potential customer / user.
21–30 of 178 posts
To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
I get running these models is not cheap, but they just lost a potential customer / user.
This is ridiculous. 32B and beating deepseek and o1. And yet I'm trying it out and, yeah, it seems pretty intelligent... Remember when models this size could just about maintain a conversation?
This is insane matching deepseek but 20x smaller?
I wonder if having a big mixture of experts isn't all that valuable for the type of tasks in math and coding benchmarks. Like my intuition is that you need all the extra experts because models store fuzzy knowledge in their feed-forward layers, and having a lot of feed-forward weights lets you store a longer tail of knowledge. Math and coding benchmarks do sometimes require highly specialized knowledge, but if we bel…
Wasn't this release in Nov 2024 as a "preview" with similarly impressive performance? https://qwenlm.github.io/blog/qwq-32b-preview/
To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.
To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.
To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.