Live data from Hacker News

QwQ-32B: Embracing the Power of Reinforcement Learning

qwenlm.github.io

131–140 of 178 posts

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#131
post #96

Earlier quoted context omitted.

Ok, this explains why QwQ is working great on their chat. Btw I saw this thing multiple times: that ollama inference, for one reason or the other, even without quantization, somewhat had issues with the actual model performance. In one instance the same model with the same quantization level, if run with MLX was great, and I got terrible results with ollama: the point here is not ollama itself, but there is no testin…

Yeah the state of the art is pretty awful. There have been multiple incidents where a model has been dropped on ollama with the wrong chat template, resulting in it seeming to work but with greatly degraded performance. And I think it's always been a user that notices, not the ollama team or the model team.

I'm grateful for anyone's contributions to anything, but I kinda shake my head about ollama. the reason stuff like this happens is they're doing the absolute minimal job necessary, to get the latest model running, not working.

I make a llama.cpp wrapper myself, and it's somewhat frustrating putting effort in for everything from big obvious UX things, like error'ing when the context is too small for your input instead of just making you think the model is crap, to long-haul engineering commitments, like integrating new models with llama.cpp's new tool calling infra, and testing them to make sure it, well, actually works.

I keep telling myself that this sort of effort pays off a year or two down the road, once all that differentiation in effort day-to-day adds up. I hope :/

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#132

Earlier quoted context omitted.

Tariffs don't create local jobs, they shut down exporting industries (other countries buy our exports with the dollars we pay them for our imports) and some of those people may over time transition to non-export industries. Here's an analysis indicating how many jobs would be destroyed in total over several scenarios: https://taxfoundation.org/research/all/federal/trump-tariffs...

They will, out of sheer necessity. Local industries will be incentivized to restart. And of course, there are already carve-outs for the automotive sector that needs steel, overseas components, etc. I expect more carve-outs will be made, esp. for the military. I don't think the tariffs are being managed intelligently, but they will have the intended effect of moving manufacturing back to the US, even if, in the short…

> even if, in the short term, it's going to inflate prices, and yes, put a lot of businesses in peril.

This is optimistic. They could totally inflate prices in the long term, and not just create inflation, but reduce the standard of living Americans are used to. That in itself is fine as Americans probably consume too much, but living in the USA will become more like living in Europe where many goods are much more expensive.

Worst case is that American Juche turns out to be just like North Korean Juche.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#134
post #39

Note the massive context length (130k tokens). Also because it would be kinda pointless to generate a long CoT without enough context to contain it and the reply. EDIT: Here we are. My first prompt created a CoT so long that it catastrophically forgot the task (but I don't believe I was near 130k -- using ollama with fp16 model). I asked one of my test questions with a coding question totally unrelated to what it say…

Can’t wait to see if my memory can even acocomodate this context

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#135
post #29

Earlier quoted context omitted.

super impressive. we won't need that many GPUs in the future if we can have the performance of DeepSeek R1 with even less parameters. NVIDIA is in trouble. We are moving towards a world of very cheap compute: https://medium.com/thoughts-on-machine-learning/a-future-of-...

Have you heard of Jevons paradox? That says that whenever new tech is used to make something more efficient the tech is just upscaled to make the product quality higher. Same here. Deepseek has some algoritmic improvements that reduces resources for the same output quality. But increasig resources (which are available) will increase the quality. There will be always need for more compute. Nvidia is not in trouble. Th…

yeah, sure, I guess the investors selling NVIDIA's stock like crazy know nothing about jevons

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#137
post #93

Earlier quoted context omitted.

Doesn't every odd number has a e ? one three five seven nine Is this a riddle which has no answer ? or what? why are people on internet saying its answer is one huh??

given one, three, five, seven, nine (odd numbers), seems like the machine should have said "there are no odd numbers without an e" since every odd number ends in an odd number, and when spelling them you always have to.. mention the final number. these LLM's don't think too well. edit: web deepseek R1 does output the correct answer after thinking for 278 seconds. The funny thing is it answered because it seemingly ga…

https://www.youtube.com/watch?v=IFcyYnUHVBA

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#138
post #31

I guess I won’t be needing that 512GB M3 Ultra after all.

A max with 64 GB of ram should be able to run this (I hope). I have to wait until an MLX model is available to really evaluate its speed, though.

Yep, it does that. I have 64 GB and was actually running 40 GB of other stuff.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#139

Earlier quoted context omitted.

Protectionist policy, if applied consistently, will actually lead to more jobs (and higher wages) eventually , but also higher inflation and job losses in the short term, and a more insular economy. It's foolish to go so hard, and so fast – or this is just a negotiation tactic – so I think the Trump administration is going to compromise by necessity, but in time supply chains will adjust to the new reality, and tarif…

That's an assumption, I'm trying to challenge it. Taxes usually take money out of the economy and lead to less activity. Why should a (very high) tax on transportation be different? These are not the sorts of things we can afford to just do without making sure they will work.

It's a debate that has been had by many people far more informed than anyone who will see this thread, many times over decades or even a few centuries. Rather than challenging it on a very basic level (it's a tax, all taxes are bad, why should this tax be different), just look up the other debates and read them.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#140
post #7

To test: https://chat.qwen.ai/ and select Qwen2.5-plus, then toggle QWQ.

They baited me into putting in a query and then asking me to sign up to submit it. Even have a "Stay Logged Out" button that I thought would bypass it, but no. I get running these models is not cheap, but they just lost a potential customer / user.

Check out venice.ai

They're pretty up to date with latest models. $20 a month

Post reply on HN