Interestingly, it failed today's NY Times Connections while 01-preview nailed it. The prompt if anyone wants to try it: • ENDEAVOR • CURB • NATIONAL • BOARDWALK • HERTZ • TWIN • MOLE • ENTERPRISE • SILICON • PROJECT • TIGER • VOLT • GAME • RAY • SECOND • VENTURE Its a game of NY Times connections. You need to make 4 groups of 4 words. Can you do it?
False start with car rental companies, only three, not four. Hertz has to be a unit then, making the first group hertz, second, mole, and volt. Similarly, enterprise must be in the business sense. Enterprise, project, venture, endeavor. Tiger, Ray, National, Twin are all singular versions of baseball teams? Curb, silicon, boardwalk, and game are left. Boardwalk is the most valuable monopoly property, silicon is the m…
QwQ: Alibaba's O1-like reasoning LLM
251–260 of 435 posts
Re: QwQ: Alibaba's O1-like reasoning LLM
#252Earlier quoted context omitted.
> everyone Let's not disrespect the team working on Qwen, these folks have shown that they are able to ship models that are better than everybody else's in the open weight category. But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point.
> But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point. That's like saying that CocaCola has no other moat than the CocaCola trademark. That's an extremely powerful moat to have indeed.
Re: QwQ: Alibaba's O1-like reasoning LLM
#253Earlier quoted context omitted.
In fairness it actually works out the correct answer fairly quickly (20 lines, including a false start and correction thereof). It seems to have identified (correctly) that this is a tricky question that it is struggling with so it does a lot of checking.
> Let me check online for similar problems. And finally googles the problem, like we do :)
Re: QwQ: Alibaba's O1-like reasoning LLM
#254Earlier quoted context omitted.
Well, the second they'll start overwhelmingly outperforming other open source LLMs, and people start incorporating them into their products, they'll get banned in the states. I'm being cynical, but the whole "dangerous tech with loads of backdoors built into it" excuse will be used to keep it away. Whether there will be some truth to it or not, that's a different question.
Qwen models have ideological backdoors already. They rewrite history, deny crimes from the regime, and push the CCP narratives. Even if their benchmarks are impressive, I refuse to ship any product with it. I'll stick with Llama and Gemma for now.
Re: QwQ: Alibaba's O1-like reasoning LLM
#255> This version is but an early step on a longer journey - a student still learning to walk the path of reasoning. Its thoughts sometimes wander, its answers aren’t always complete, and its wisdom is still growing. But isn’t that the beauty of true learning? To be both capable and humble, knowledgeable yet always questioning? > Through deep exploration and countless trials, we discovered something profound: when given…
how much are you willing to bet that it was written by a human
Re: QwQ: Alibaba's O1-like reasoning LLM
#256This one is pretty impressive. I'm running it on my Mac via Ollama - only a 20GB download, tokens spit out pretty fast and my initial prompts have shown some good results. Notes here: https://simonwillison.net/2024/Nov/27/qwq/
It simply did not want to use XML tools for some reason something that even qwen coder does not struggle with: https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...
I have not seen any model including sonnet that is able to 1 shot a working 9x9 go board
For ref gpt-4o which is still quite bad https://discuss.samsaffron.com/discourse-ai/ai-bot/shared-ai...
Re: QwQ: Alibaba's O1-like reasoning LLM
#257Re: QwQ: Alibaba's O1-like reasoning LLM
#258Re: QwQ: Alibaba's O1-like reasoning LLM
#259Earlier quoted context omitted.
Well, the second they'll start overwhelmingly outperforming other open source LLMs, and people start incorporating them into their products, they'll get banned in the states. I'm being cynical, but the whole "dangerous tech with loads of backdoors built into it" excuse will be used to keep it away. Whether there will be some truth to it or not, that's a different question.
Qwen models have ideological backdoors already. They rewrite history, deny crimes from the regime, and push the CCP narratives. Even if their benchmarks are impressive, I refuse to ship any product with it. I'll stick with Llama and Gemma for now.
I can't comment on the particular one, but I feel like this will unfortunately apply to most works out of authoritarian regimes. As a researcher/organization living under strict rule that can be oppressive, do you really risk releasing something that would get you into trouble? A model that would critique the government or acknowledge events they'd rather pretend don't exist? Actually, if not for the financial possibilities, working with LLMs in general could open one up to some pretty big risks, if the powers that be don't care about the inherent randomness of the technology.
Re: QwQ: Alibaba's O1-like reasoning LLM
#260Earlier quoted context omitted.
> But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point. That's like saying that CocaCola has no other moat than the CocaCola trademark. That's an extremely powerful moat to have indeed.
It's not nothing, but it's very different from the CocaCola case. Coke is built on getting people used to the specific taste, especially when they're young, and taste is an almost subconscious/System1 thing where people seek comfort in the familiar. ChatGPT is more of a System2 thing, people don't really care whether the answer they seek comes from ChatGPT or from AskJeeves, they just want answers of reasonably simil…