Live data from Hacker News

QwQ-32B: Embracing the Power of Reinforcement Learning

qwenlm.github.io

141–150 of 178 posts

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#141

Earlier quoted context omitted.

Yeah the state of the art is pretty awful. There have been multiple incidents where a model has been dropped on ollama with the wrong chat template, resulting in it seeming to work but with greatly degraded performance. And I think it's always been a user that notices, not the ollama team or the model team.

I'm grateful for anyone's contributions to anything, but I kinda shake my head about ollama. the reason stuff like this happens is they're doing the absolute minimal job necessary, to get the latest model running , not working. I make a llama.cpp wrapper myself, and it's somewhat frustrating putting effort in for everything from big obvious UX things, like error'ing when the context is too small for your input instea…

Can you link your wrapper? I've read and run up against a lot of footguns related to Ollama myself and I think surfacing community efforts to do better would be quite useful.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#142
post #80

Told it to generate a Handbrake CLI command for some specific transcoding requirements, it thought for 30+ seconds and produced only CoT, no output. Needs work, lol.

Check your context settings on ollama if that's what you're using to run it and override the proper environment variables. By default, its 2048 iirc.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#143
post #33

Earlier quoted context omitted.

Is this the best way to run your own models these days?

It's the easiest to setup, but you can get 2x-6x faster with TGI and vLLM depending on the scenario.

vllm isn't even hard to setup!

I find it so funny that HN is sitting in the stoneage with LLM inference.

Meanwhile I'm here with sillytavern hooked to my own vllm server, getting crazy fast performance on my models and having a complete suite of tools for using LLMs.

Most folks on here have never heard of sillytavern, or oobabooga, or any of the other projects for LLM UI/UX (LM-studio). It's insanity that there hasn't been someone like ADOBE building a pro/prosumer UI for LLMs yet.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#144

Earlier quoted context omitted.

I'm grateful for anyone's contributions to anything, but I kinda shake my head about ollama. the reason stuff like this happens is they're doing the absolute minimal job necessary, to get the latest model running , not working. I make a llama.cpp wrapper myself, and it's somewhat frustrating putting effort in for everything from big obvious UX things, like error'ing when the context is too small for your input instea…

Can you link your wrapper? I've read and run up against a lot of footguns related to Ollama myself and I think surfacing community efforts to do better would be quite useful.

Cheers, thanks for your interest:

Telosnex, @ telosnex.com --- fwiw, general positioning is around paid AIs, but there's a labor-of-love llama.cpp backed on device LLM integration that makes them true peers, both in UI and functionality. albeit with a warning sign because normie testers all too often wander into trying it on their phone and killing their battery.

My curse is the standard engineer one - only place I really mention it is one-off in comments like here to provide some authority on a point I want to make...I'm always one release away from it being perfect enough to talk up regularly.

I really really need to snap myself awake and ban myself from the IDE for a month.

But this next release is a BFD, full agentic coding, with tons of tools baked in, and I'm so damn proud to see the extra month I've spent getting llama.cpp tools working agentically too. (https://x.com/jpohhhh/status/1897717300330926109, real thanks is due to @ochafik at Google, he spent a very long term making a lot of haphazard stuff in llama.cpp coalesce. also phi-4 mini. this is the first local LLM that is reasonably fast and an actual drop-in replacement for RAG and tools, after my llama.cpp patch)

Please, feel free to reach out if you try it and have any thoughts, positive or negative. james @ the app name.com

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#145

Earlier quoted context omitted.

If I had to guess, more tariffs and sanctions that increase the competing nation's self-reliance and harm domestic consumers. Perhaps my peabrain just can't comprehend the wisdom of policymakers on the sanctions front, but it just seems like all it does is empower the target long-term.

The tarrifs are for the US to build it's own domestic capabilities, but this will ultimately shift the rest of the world's trade away from the US and toward each other. It's a trade-off – no pun intended – between local jobs/national security and downgrading their own economy/geo-political standing/currency. Anyone who's been making financial bets on business as usual for globalization is going to see a bit of a spee…

The tariffs are seen as "free money" that will allow for cutting taxes on the wealthy. Note that the current messaging is "we spend too much money" and there's nothing about "we need to invest in _foo_"

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#146

Dude its so crazy , in my personal experience , I gave it can you read what I have wrote backwards and answer that query ip fo eulav si tahw profile Qwen2.5-Max 11:22 am Thinking completed Okay, let me try to figure this out. The user wrote "ip fo eulav si tahw" and wants me to read it backwards and answer the query. Hmm, first, I need to reverse the entire string. Let's see, reversing "ip fo eulav si tahw" would be…

The example you gave is not very impressive, normal, non-reasoning LLMs have been able to do this for a while. E.g., Claude 3.5 Haiku solves this no problem.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#147

Earlier quoted context omitted.

You can also set OLLAMA_CONTEXT_LENGTH= as an environment variable to change ollama's default context length.

I tried this with ollama run, and it had no effect at all.

that env parameter is brand new, did you update ollama?

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#148
post #105

Earlier quoted context omitted.

From https://huggingface.co/Qwen/QwQ-32B Presently, vLLM only supports static YARN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts. We advise adding the rope_scaling configuration only when processing long contexts is required.

Sorry, could you please explain what this means? I'm not into machine learning, so I don't get the jargon.

Well I can't be positive, but it looks like some of the factors that support a long context length might be set wrong. https://blog.eleuther.ai/yarn/

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#149

Earlier quoted context omitted.

They will, out of sheer necessity. Local industries will be incentivized to restart. And of course, there are already carve-outs for the automotive sector that needs steel, overseas components, etc. I expect more carve-outs will be made, esp. for the military. I don't think the tariffs are being managed intelligently, but they will have the intended effect of moving manufacturing back to the US, even if, in the short…

> even if, in the short term, it's going to inflate prices, and yes, put a lot of businesses in peril. This is optimistic. They could totally inflate prices in the long term, and not just create inflation, but reduce the standard of living Americans are used to. That in itself is fine as Americans probably consume too much, but living in the USA will become more like living in Europe where many goods are much more ex…

> They could totally inflate prices in the long term, and not just create inflation

This will all happen. But as I said, this is a trade-off. Devalue the currency, incentivize local production, increase exports, revive the working class – that's the long term goal.

> but reduce the standard of living Americans are used to.

Whose standard of living though? It's well and good if you're in a comfy desk job with health care and a pension. The discontent that led to Trump's rise is real, and it's routinely overlooked when considering how to counter him. Of the everyday people, those who have stable jobs and purpose aren't voting for Trump. (Of the wealthy, it's probably a lot more cynical who voted for Trump)

I'm not in favor of the policy, the manner in which it's being applied, or the people that are doing it, but reversing off-shoring is a consequence of using protectionist policies – be it tariffs, or subsidies.

High-skill work, and pencil-pushing desk jobs don't cover 100% of the population, and has lead to a lot of unproductive busy-work in the cities. The offshoring of blue-collar work bred the discontent that led to Trump. Trump fancies himself the new William McKinley and is using the cudgel of tariffs to re-onshore manufacturing. This is a process he started in his first administration, that was retained by Biden, and now he's doubling down and doing exactly what he promised he would – and somehow his voters are surprised?

Worse still, those service economy jobs keeping the coastal cities alive (both low skill and high skill) are on the verge of being replaced by AI – whether that's one year or 20, I don't know—though I'm wagering the latter. Physical labor is going to become more valuable as robotics is still way behind in technological development. I don't have a crystal ball, but I'd wager that–at least counterfactually—the US will have more jobs by enacting protectionist policy.

> Worst case is that American Juche turns out to be just like North Korean Juche.

Do you really in your heart of hearts think this is going to happen? I'm pretty sure the subjugation of the American people by the government would be feasible, let alone easy.

Re: QwQ-32B: Embracing the Power of Reinforcement Learning

#150
post #118

Earlier quoted context omitted.

Protectionist policy, if applied consistently, will actually lead to more jobs (and higher wages) eventually , but also higher inflation and job losses in the short term, and a more insular economy. It's foolish to go so hard, and so fast – or this is just a negotiation tactic – so I think the Trump administration is going to compromise by necessity, but in time supply chains will adjust to the new reality, and tarif…

> and higher wages) eventually Higher real wages? Do gains from trade not exist? Comparative advantage: Country A has an easier time making X than Y, and country B has an easier time making Y than X, so country A should trade some of their Xs for Ys, and both countries end up richer. I think there's some reasons to dial back interdependence a little, but I don't think it's a path likely to lead to greater wealth or r…

I don't believe I have to point this out, but this is not a policy that I think is good, it's just one that will have the intended affect of onshoring manufacturing jobs. And I'm not talking about higher wages for quants, or MBAs, or HR, or software evangelists, or door-to-door salespeople, or cashiers at Dollar General, I'm talking about for people who are underemployed, unemployed, or doing some nonsense busywork because the manufacturing sector has been eroded away over the last 4 decades.

> And certainly no reason to make erratic changes at large scale, focusing on allies and neighbors first

Those people who benefited from globalization, and who didn't care about the working class, are exactly who brought us to this moment. And I have a huge shrug to those who are loath to accept that. If only it was attended to sooner by a more sensible administration.

Post reply on HN