Live data from Hacker News

QwQ: Alibaba's O1-like reasoning LLM

qwenlm.github.io

351–360 of 435 posts

Re: QwQ: Alibaba's O1-like reasoning LLM

#351

Earlier quoted context omitted.

Macs that can run it are quite a bit more expensive than a 3090. GPUs can also do finetuning and run other models with larger batch sizes which Macs would struggle with. Also, for the models that fit both, an nvidia card can run it much faster.

I see often MBP with 48-64gb and 1TB under than 3500 CHF. Including the M4 (thx to black Friday week). Meanwhile 4090 are close to 2000CHF. I have no doubt where the actual value is.

Alright but being Swiss you can probably afford to buy both and still have most of your monthly pay left over lmao.

Re: QwQ: Alibaba's O1-like reasoning LLM

#352
post #275

Earlier quoted context omitted.

Switched to the wrong Qube that's logged into your alt just now. :) Maybe that kind of opsec failure took place earlier too.

Why do you think it's not intentional? I just replied on my phone in the elevator while going home. The other device is home laptop I share with wife. Don't need opsec in my living room :) Anyhow, you can test my findings yourself, I told you details of my prompts. Why do you think Chinese are not censoring?

They are probably censoring. It is too hard to fight the temptation of playing with the weights since nobody would know.

Re: QwQ: Alibaba's O1-like reasoning LLM

#353

It still fails at very simple stuff. E.g. "I put an ordinary rock into a glass of water. I then turn the glass of water upside down, do a little dance, and then turn the glass right side up again. Where is the rock now?" 100+ lines later... "The rock is at the bottom of the glass, submerged in the water." Models from a year ago get this right sometimes https://pastebin.com/em5TT4Zn

[deleted]

Re: QwQ: Alibaba's O1-like reasoning LLM

#354
post #58

Earlier quoted context omitted.

I don’t see why they wouldn’t. If you’re China and willing to pour state resources into LLMs, it’s an incredible ROI if they’re adopted. LLMs are black boxes, can be fine tuned to subtly bias responses, censor, or rewrite history. They’re a propaganda dream. No code to point to of obvious interference.

That is a pretty dark view on almost 1/5th of humanity and a nation with a track record of giving the world important innovations: paper making, silk, porcelain, gunpowder and compass to name the few. Not everything has to be around politics.

Nation/culture != the current regime

Re: QwQ: Alibaba's O1-like reasoning LLM

#355

We are lucky that Alibaba, Meta and Mistral sees some strategic value in public releases. If we it was just one of them, it would be a fragile situation for downstream startups. And they’re even situated in three different countries.

They're salting the earth / destroying the moat for OpenAI and Anthropic, and to a lesser extent, for Google. Basically, now pretty much anyone can do "sufficient" AI right in their garage. I use Mistral Large, and most of the time its output is easily on par with GPT4, with total privacy. This destruction of the moat prevents the future where e.g. OpenAI becomes your interface to the internet, recommendations, shopping, social media, news, etc. The moment you start doing something that's clearly valuable, 10 other companies pop up and destroy your potential future margin. They are trying the "internet search" part of that, because that needs a search index, and that's not something that's easy to do. We're lucky that Google blew its AI lead so badly - it _could_ realistically become all of the above, if it wasn't so badly mismanaged.

Re: QwQ: Alibaba's O1-like reasoning LLM

#357
Tried this yesterday under Ollama. Asked it to explain and implement bitonic sort. It churned for a little bit, and started putting together an implementation towards the end, but then seemingly ran out of generation window because the CoT was so long winded. TL;DR: promising, but needs more work. I'm sure they'll improve it greatly in the months to come, the potential is pretty clearly there.

Re: QwQ: Alibaba's O1-like reasoning LLM

#358

Earlier quoted context omitted.

Macs that can run it are quite a bit more expensive than a 3090. GPUs can also do finetuning and run other models with larger batch sizes which Macs would struggle with. Also, for the models that fit both, an nvidia card can run it much faster.

I see often MBP with 48-64gb and 1TB under than 3500 CHF. Including the M4 (thx to black Friday week). Meanwhile 4090 are close to 2000CHF. I have no doubt where the actual value is.

A 3090 often goes for $800-900 on the used market in the US. Two of these would be $1800, and you get a much more versatile machine. However, the downside is also obvious, since your two 3090s can draw up to 800 W, and there's no chance you can carry it around with you. Overall, it's not that obvious where the actual value is, as it depends a lot on what you want to use it for.

Re: QwQ: Alibaba's O1-like reasoning LLM

#359

We are lucky that Alibaba, Meta and Mistral sees some strategic value in public releases. If we it was just one of them, it would be a fragile situation for downstream startups. And they’re even situated in three different countries.

Maybe 'public AI' is a better term than 'open' ai

Re: QwQ: Alibaba's O1-like reasoning LLM

#360
post #332

Earlier quoted context omitted.

> We are lucky that Alibaba, Meta and Mistral sees some strategic value in public releases. Now if we only can get Meta to understand what "Open Source" means so the word doesn't lose all meaning in the future.

Words mean whatever's convenient to the bottom line, which is why the OSI (a consortium of Amazon, Google, Microsoft etc) still doesn't recognize the SSPL, as it would be particularly inconvenient for clouds.

So you're saying open source doesn't exist to be free labor for SaaS?

The OSI is fully captured by companies with a vested interest in promoting that model and/or using open source to 'dump' on the market and commoditize their compliments. To recapture the spirit of open source as being about freedom for actual users (as opposed to free labor for jailed SaaS) and a mutualistic gift culture (as opposed to a take-take-take culture) probably requires abandoning the OSI.

Post reply on HN