QwQ-32B: Embracing the Power of Reinforcement Learning
31–40 of 178 posts
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#32Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#33Available on ollama now as well.
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#34To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
super impressive. we won't need that many GPUs in the future if we can have the performance of DeepSeek R1 with even less parameters. NVIDIA is in trouble. We are moving towards a world of very cheap compute: https://medium.com/thoughts-on-machine-learning/a-future-of-...
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#35I guess I won’t be needing that 512GB M3 Ultra after all.
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#36To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
super impressive. we won't need that many GPUs in the future if we can have the performance of DeepSeek R1 with even less parameters. NVIDIA is in trouble. We are moving towards a world of very cheap compute: https://medium.com/thoughts-on-machine-learning/a-future-of-...
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#37To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#38Chinese strategy is open-source software part and earn on robotics part. And, They are already ahead of everyone in that game. These things are pretty interesting as they are developing. What US will do to retain its power? BTW I am Indian and we are not even in the race as country. :(
Re: QwQ-32B: Embracing the Power of Reinforcement Learning
#39EDIT: Here we are. My first prompt created a CoT so long that it catastrophically forgot the task (but I don't believe I was near 130k -- using ollama with fp16 model). I asked one of my test questions with a coding question totally unrelated to what it says:
But the problem is in this question. Wait perhaps I'm getting ahead of myself.
Wait the user hasn't actually provided a specific task yet. Let me check again.
The initial instruction says:
"Please act as an AI agent that can perform tasks... When responding, first output a YAML data structure with your proposed action, then wait for feedback before proceeding."
But perhaps this is part of a system prompt? Wait the user input here seems to be just "You will be given a problem. Please reason step by step..." followed by a possible task?
Note: Ollama "/show info" shows that the context size set is correct.