Earlier quoted context omitted.
What hardware are you able to run this on?
If your job or hobby in any way likes LLMs, and you like to "Work Anywhere", it's hard not to justify the MBP Max (e.g. M3 Max, now M4 Max) with 128GB. You can run more than you'd think, faster than you'd think. See also Hugging Face's MLX community: https://huggingface.co/mlx-community QwQ 32B is featured: https://huggingface.co/collections/mlx-community/qwq-32b-pre... If you want a traditional GUI, LM Studio beta 0…
QwQ: Alibaba's O1-like reasoning LLM
261–270 of 435 posts
Re: QwQ: Alibaba's O1-like reasoning LLM
#262Earlier quoted context omitted.
I don’t see why they wouldn’t. If you’re China and willing to pour state resources into LLMs, it’s an incredible ROI if they’re adopted. LLMs are black boxes, can be fine tuned to subtly bias responses, censor, or rewrite history. They’re a propaganda dream. No code to point to of obvious interference.
That is a pretty dark view on almost 1/5th of humanity and a nation with a track record of giving the world important innovations: paper making, silk, porcelain, gunpowder and compass to name the few. Not everything has to be around politics.
Re: QwQ: Alibaba's O1-like reasoning LLM
#263Earlier quoted context omitted.
In fairness it actually works out the correct answer fairly quickly (20 lines, including a false start and correction thereof). It seems to have identified (correctly) that this is a tricky question that it is struggling with so it does a lot of checking.
> Let me check online for similar problems. And finally googles the problem, like we do :)
Re: QwQ: Alibaba's O1-like reasoning LLM
#264This one is crazy. I made up a silly topology problem which I guessed wouldn't be in a textbook (given X create a shape with Euler characteristic X) and set it to work. Its first effort was a program that randomly generated shapes, calculated X and hoped it was right. I went and figured out a solution and gave it a clue. Watching it "think" through the answer is surreal and something I haven't felt since watching GPT…
Re: QwQ: Alibaba's O1-like reasoning LLM
#265Re: QwQ: Alibaba's O1-like reasoning LLM
#266It gets the Sally question correct, but it takes more than 100 lines of reasoning. >Sally has three brothers. Each brother has two sisters. How many sisters does sally have? Here is the answer: https://pastebin.com/JP2V92Kh
"But that doesn't make sense because Sally can't be her own sister."
Having said this, how many 'lines' of reasoning does the average human need? It's a weird comparison perhaps but the point is does it really matter if it needs 100 or 100k 'lines', if it could hide that (just as we hide our thoughts or even can't really access the - semi-parallel - things our brain does to come to an answer) eventually and summarise it + give the correct answer, that'd be acceptable?
Re: QwQ: Alibaba's O1-like reasoning LLM
#267This one is crazy. I made up a silly topology problem which I guessed wouldn't be in a textbook (given X create a shape with Euler characteristic X) and set it to work. Its first effort was a program that randomly generated shapes, calculated X and hoped it was right. I went and figured out a solution and gave it a clue. Watching it "think" through the answer is surreal and something I haven't felt since watching GPT…
I mean, isn't it trivial? For instance, you can make any Euler characteristic you want by making a cell complex with a 1-cells and b 2-cells and then the Euler characteristic will be 1 + b - a. Am I missing something? If this is the problem and this is how the model looks for a solution, yes, it is overcomplicating things.
Yes that's a nice solution. I think I asked for the answer as a polyhedron (I tried a few variations), but maybe it (and I) should have just reached for CW complexes.
Re: QwQ: Alibaba's O1-like reasoning LLM
#268I asked the classic 'How many of the letter “r” are there in strawberry?' and I got an almost never ending stream of second guesses. The correct answer was ultimately provided but I burned probably 100x more clockcycles than needed. See the response here: https://pastecode.io/s/6uyjstrt
Well, to be perfectly honest, it's hard question for an LLM that reasons in tokens and not letters. Reminds me of that classic test that kids easily pass and grownups utterly fail. The test looks like this: continue a sequence: 0 - 1 5 - 0 6 - 1 7 - 0 8 - 2 9 - ? Grownups try to find a pattern in the numbers, different types of series, progressions, etc. The correct answer is 1 because it's the number of circles in t…
[...]
Starting with 0 paired with 1. If I think about functions or some formula that could map these values, but that might be too complicated.
Alternatively, perhaps these numbers represent something specific. For example, maybe they're representing the number of certain features in the number itself, like the number of closed loops in the digit's shape.
Let's think about that:
0 has one closed loop. 5 has no closed loops. 6 has one closed loop. 7 has no closed loops. 8 has two closed loops. 9 has one closed loop. Wait a minute, that seems to match the sequence:
0 (1 loop) - 1
5 (0 loops) - 0
6 (1 loop) - 1
7 (0 loops) - 0
8 (2 loops) - 2
So, if 9 has one closed loop, then it should be paired with 1.
But let's double-check this theory with the given pairs.
0 has one loop, paired with 1 – matches.
5 has no loops, paired with 0 – matches.
6 has one loop, paired with 1 – matches.
7 has no loops, paired with 0 – matches.
8 has two loops, paired with 2 – matches.
Yes, this seems consistent.
Therefore, 9 should have one closed loop, so it should be paired with 1.
[...]
Then it continued trying another approach related to the number of angles in the digit shape, but it discarded it. Then "Alternatively, perhaps it's related to the number of segments that are 'on' in a seven-segment display." and so on.
Re: QwQ: Alibaba's O1-like reasoning LLM
#269Earlier quoted context omitted.
What hardware are you able to run this on?
If your job or hobby in any way likes LLMs, and you like to "Work Anywhere", it's hard not to justify the MBP Max (e.g. M3 Max, now M4 Max) with 128GB. You can run more than you'd think, faster than you'd think. See also Hugging Face's MLX community: https://huggingface.co/mlx-community QwQ 32B is featured: https://huggingface.co/collections/mlx-community/qwq-32b-pre... If you want a traditional GUI, LM Studio beta 0…
I've been off Mac's for ten years since OSX started driving me crazy, but I've been strongly considering picking up the latest Mac Mini as a poor man's version of what you're talking about. For €1k you can get an M4 with 32GiB of unified ram, of an M4 pro with 64GiB for €2k which is a bit more affordable.
If you shucked the cheap ones into your rack you could have a very hefty little Beowulf cluster for the price of that MBP.
Re: QwQ: Alibaba's O1-like reasoning LLM
#270You must use math questions that have never entered the training data set for testing to know whether LLM has real reasoning capabilities. https://venturebeat.com/ai/ais-math-problem-frontiermath-ben...
I like the tentativeness, I see a lot of : wait, But, perhaps, maybe, This is getting too messy, this is confusing, that can't be right, this is getting too tricky for me right now, this is very difficult.
I kind of find it harder to not anthropomorphise when comparing with ChatGPT. It feels like it's trying to solve it from first principles but with the depth of Highschool Physics knowledge.