Live data from Hacker News

QwQ: Alibaba's O1-like reasoning LLM

qwenlm.github.io

101–110 of 435 posts

Re: QwQ: Alibaba's O1-like reasoning LLM

#101

Earlier quoted context omitted.

> everyone Let's not disrespect the team working on Qwen, these folks have shown that they are able to ship models that are better than everybody else's in the open weight category. But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point.

They have the moat of being able to raise large funding rounds than everybody else: Access to capital.

Do they have more access to capital than the CCP, if the latter decided to put its efforts behind Alibaba on this? Genuine question.

Re: QwQ: Alibaba's O1-like reasoning LLM

#102
post #94

Earlier quoted context omitted.

Of all the types of tokens in the world video is not the one that comes to mind as having a shortage. By setting a a few thousand security cameras in various high traffic places you can get almost infinite footage. Instagram, Youtube and Snapchat have no shortage of data too.

except 1) tiktok is video stream data many orders of magnitude larger than any security cam data, that's attached to real identity 2) china doesn't have direct access to Instagram reels and shorts, so yeah

Why does tying it to identity help LLM training?

It's pretty unclear that having orders of magnitude more video data of dancing is useful. Diverse data is much useful!

Re: QwQ: Alibaba's O1-like reasoning LLM

#103
post #84

> Find the least odd prime factor of 2019^8+1 God that's absurd. The mathematical skills involved on that reasoning are very advanced; the whole process is a bit long but that's impressive for a model that can potentially be self-hosted.

The process is only long because it babbled several useless ideas (direct factoring, direct exponentiating, Sophie Germain) before (and in the middle of) the short correct process.

I think it's exploring in-context. Bringing up related ideas and not getting confused by them is pivotal to these models eventually being able to contribute as productive reasoners. These traces will be immediately helpful in a real world iterative loop where you don't already know the answers or how to correctly phrase the questions.

Re: QwQ: Alibaba's O1-like reasoning LLM

#104

Earlier quoted context omitted.

Well, the second they'll start overwhelmingly outperforming other open source LLMs, and people start incorporating them into their products, they'll get banned in the states. I'm being cynical, but the whole "dangerous tech with loads of backdoors built into it" excuse will be used to keep it away. Whether there will be some truth to it or not, that's a different question.

[flagged]

You are absolutely correct. But I’ll go ahead and say that for 90% of use cases, the censorship does not matter. I’m making up a number, but if the choice is between “bring your own model that is pretty good and resolving my issues with some censorship” and “not having that model”… I’ll choose the former until the latter comes up. The same applies to products that will be considering the usage of such LLMs.

Re: QwQ: Alibaba's O1-like reasoning LLM

#105

Seems that given enough compute everyone can build a near-SOTA LLM. So what is this craze about securing AI dominance?

> everyone Let's not disrespect the team working on Qwen, these folks have shown that they are able to ship models that are better than everybody else's in the open weight category. But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point.

> But fundamentally yes, OpenAI has no other moat than the ChatGPT trademark at this point.

That's like saying that CocaCola has no other moat than the CocaCola trademark.

That's an extremely powerful moat to have indeed.

Re: QwQ: Alibaba's O1-like reasoning LLM

#106
post #74

This one is pretty impressive. I'm running it on my Mac via Ollama - only a 20GB download, tokens spit out pretty fast and my initial prompts have shown some good results. Notes here: https://simonwillison.net/2024/Nov/27/qwq/

What hardware are you able to run this on?

Re: QwQ: Alibaba's O1-like reasoning LLM

#107
post #95
post #86

Earlier quoted context omitted.

This. I'm 100% certain that Chinese models are not long for this market. Whether or not they are free is irrelevant. I just can't see the US government allowing us access to those technologies long term.

I disagree, that is really only police-able for online services. For local apps, which will eventually include games, assistants and machine symbiosis, I expect a bring your own model approach.

How many people do you think will ever use “bring your own model” approach? Those numbers are so statistically insignificant that nobody will bother when it comes to making money. I’m sure we will hack our way through it, but if it’s not available to general public, those Chinese companies won’t see much market share in the west.

Re: QwQ: Alibaba's O1-like reasoning LLM

#108

Seems that given enough compute everyone can build a near-SOTA LLM. So what is this craze about securing AI dominance?

AI dominance is secured through legal and regulatory means, not technical methods. So for instance, a basic strategy is to rapidly develop AI and then say “Oh wow AI is very dangerous we need to regulate companies and define laws around scraping data” and then make it very difficult for new players to enter the market. When a moat can’t be created, you resort to ladder kicking.

Operation Chokepoint 2.0

Relevant https://x.com/benaverbook/status/1861511171951542552

Re: QwQ: Alibaba's O1-like reasoning LLM

#109
post #3
post #2

Model weights and demo on HF https://huggingface.co/collections/Qwen/qwq-674762b79b75eac0...

For some fun - put in "Let's play Wordle" It seems to blabber to itself infinitely ...

It seemed to get stuck in a loop for a while for me but eventually decided "EARTY" was the solution: https://pastebin.com/VwvRaqYK
Post reply on HN