Live data from Hacker News

QwQ: Alibaba's O1-like reasoning LLM

qwenlm.github.io

251–260 of 435 posts

Re: QwQ: Alibaba's O1-like reasoning LLM

#251

Interestingly, it failed today's NY Times Connections while 01-preview nailed it. The prompt if anyone wants to try it: • ENDEAVOR • CURB • NATIONAL • BOARDWALK • HERTZ • TWIN • MOLE • ENTERPRISE • SILICON • PROJECT • TIGER • VOLT • GAME • RAY • SECOND • VENTURE Its a game of NY Times connections. You need to make 4 groups of 4 words. Can you do it?

False start with car rental companies, only three, not four. Hertz has to be a unit then, making the first group hertz, second, mole, and volt. Similarly, enterprise must be in the business sense. Enterprise, project, venture, endeavor. Tiger, Ray, National, Twin are all singular versions of baseball teams? Curb, silicon, boardwalk, and game are left. Boardwalk is the most valuable monopoly property, silicon is the m…

Seems like a very American riddle, three-quarters of these are based on assuming US-centric associations to the words. Not very surprising that a non-US-based model doesn't get there as easily as a US-based one.

Re: QwQ: Alibaba's O1-like reasoning LLM

#252

Earlier quoted context omitted.

> everyone Let's not disrespect the team working on Qwen, these folks have shown that they are able to ship models that are better than everybody else's in the open weight category. But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point.

> But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point. That's like saying that CocaCola has no other moat than the CocaCola trademark. That's an extremely powerful moat to have indeed.

It's not nothing, but it's very different from the CocaCola case. Coke is built on getting people used to the specific taste, especially when they're young, and taste is an almost subconscious/System1 thing where people seek comfort in the familiar. ChatGPT is more of a System2 thing, people don't really care whether the answer they seek comes from ChatGPT or from AskJeeves, they just want answers of reasonably similar quality. There's still value in branding and name recognition, but the switching costs are much lower here.

Re: QwQ: Alibaba's O1-like reasoning LLM

#253
post #238

Earlier quoted context omitted.

In fairness it actually works out the correct answer fairly quickly (20 lines, including a false start and correction thereof). It seems to have identified (correctly) that this is a tricky question that it is struggling with so it does a lot of checking.

> Let me check online for similar problems. And finally googles the problem, like we do :)

brilliant

Re: QwQ: Alibaba's O1-like reasoning LLM

#254

Earlier quoted context omitted.

Well, the second they'll start overwhelmingly outperforming other open source LLMs, and people start incorporating them into their products, they'll get banned in the states. I'm being cynical, but the whole "dangerous tech with loads of backdoors built into it" excuse will be used to keep it away. Whether there will be some truth to it or not, that's a different question.

Qwen models have ideological backdoors already. They rewrite history, deny crimes from the regime, and push the CCP narratives. Even if their benchmarks are impressive, I refuse to ship any product with it. I'll stick with Llama and Gemma for now.

If you carefully study the so-called regime oppression, you will find that in the end, nothing happened, and there were no large-scale deaths. But it was massively exaggerated by CNN and BBC only because of the appearance of weapons.

Re: QwQ: Alibaba's O1-like reasoning LLM

#255
post #232

> This version is but an early step on a longer journey - a student still learning to walk the path of reasoning. Its thoughts sometimes wander, its answers aren’t always complete, and its wisdom is still growing. But isn’t that the beauty of true learning? To be both capable and humble, knowledgeable yet always questioning? > Through deep exploration and countless trials, we discovered something profound: when given…

how much are you willing to bet that it was written by a human

Not saying it wasn't written by AI, but it also looks a lot like Chinese to English translation

Re: QwQ: Alibaba's O1-like reasoning LLM

#256
post #74

This one is pretty impressive. I'm running it on my Mac via Ollama - only a 20GB download, tokens spit out pretty fast and my initial prompts have shown some good results. Notes here: https://simonwillison.net/2024/Nov/27/qwq/

I find it odd that is refused me so badly https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai... my guess is that I am using a quantized model

It simply did not want to use XML tools for some reason something that even qwen coder does not struggle with: https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...

I have not seen any model including sonnet that is able to 1 shot a working 9x9 go board

For ref gpt-4o which is still quite bad https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...

Re: QwQ: Alibaba's O1-like reasoning LLM

#258

Earlier quoted context omitted.

[flagged]

Since this is a local model, you can trivially force it to do pretty much whatever you want by forcing the response to start with "Yes, sir!".

Any prompt or system setup examples which work well?

Re: QwQ: Alibaba's O1-like reasoning LLM

#259

Earlier quoted context omitted.

Well, the second they'll start overwhelmingly outperforming other open source LLMs, and people start incorporating them into their products, they'll get banned in the states. I'm being cynical, but the whole "dangerous tech with loads of backdoors built into it" excuse will be used to keep it away. Whether there will be some truth to it or not, that's a different question.

Qwen models have ideological backdoors already. They rewrite history, deny crimes from the regime, and push the CCP narratives. Even if their benchmarks are impressive, I refuse to ship any product with it. I'll stick with Llama and Gemma for now.

> Qwen models have ideological backdoors already. They rewrite history, deny crimes from the regime, and push the CCP narratives.

I can't comment on the particular one, but I feel like this will unfortunately apply to most works out of authoritarian regimes. As a researcher/organization living under strict rule that can be oppressive, do you really risk releasing something that would get you into trouble? A model that would critique the government or acknowledge events they'd rather pretend don't exist? Actually, if not for the financial possibilities, working with LLMs in general could open one up to some pretty big risks, if the powers that be don't care about the inherent randomness of the technology.

Re: QwQ: Alibaba's O1-like reasoning LLM

#260

Earlier quoted context omitted.

> But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point. That's like saying that CocaCola has no other moat than the CocaCola trademark. That's an extremely powerful moat to have indeed.

It's not nothing, but it's very different from the CocaCola case. Coke is built on getting people used to the specific taste, especially when they're young, and taste is an almost subconscious/System1 thing where people seek comfort in the familiar. ChatGPT is more of a System2 thing, people don't really care whether the answer they seek comes from ChatGPT or from AskJeeves, they just want answers of reasonably simil…

People can't tell coke apart from Pepsi in blind tests though, the perceived difference when people are getting asked “Is Pepsi ok?” is 100% conditioned by the attachment to the brand.
Post reply on HN