Live data from Hacker News

Qwen3: Think deeper, act faster

qwenlm.github.io

161–170 of 412 posts

Re: Qwen3: Think deeper, act faster

#162

Gotta love how Claude is always conventiently left out of all of these benchmark lists. Anthropic really is in a league of their own right now.

I think it's actually due to the fact that Claude isn't available on China, so they wouldn't be able to (legally) replicate how they evaluated the other LLMs (assuming that they didn't just use the numbers reported by each model provider)

Re: Qwen3: Think deeper, act faster

#164
post #62

Earlier quoted context omitted.

The MoE version with 3b active parameters will run significantly faster (tokens/second) on the same hardware, by about an order of magnitude (i.e. ~4t/s vs ~40t/s)

> The MoE version with 3b active parameters ~34 tok/s on a Radeon RX 7900 XTX under today's Debian 13.

And vmem use?

Re: Qwen3: Think deeper, act faster

#165
post #98

I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.

I similarly have a small, simple spatial reasoning problem that only reasoning models get right, and not all of them, and which Qwen3 on max reasoning still gets wrong. > I put a coin in a cup and slam it upside-down on a glass table. I can't see the coin because the cup is over it. I slide a mirror under the table and see heads. What will I see if I take the cup (and the mirror) away?

Yup, it flunked that one.

I also have a question that LLMs always got wrong until ChatGPT o3, and even then it has a hard time (I just tried it again and it needed to run code to work it out). Qwen3 failed, and every time I asked it to look again at its solution it would notice the error and try to solve it again, failing again:

> A man wants to cross a river, and he has a cabbage, a goat, a wolf and a lion. If he leaves the goat alone with the cabbage, the goat will eat it. If he leaves the wolf with the goat, the wolf will eat it. And if he leaves the lion with either the wolf or the goat, the lion will eat them. How can he cross the river?

I gave it a ton of opportunities to notice that the puzzle is unsolvable (with the assumption, which it makes, that this is a standard one-passenger puzzle, but if it had pointed out that I didn't say that I would also have been happy). I kept trying to get it to notice that it failed again and again in the same way and asking it to step back and think about the big picture, and each time it would confidently start again trying to solve it. Eventually I ran out of free messages.

Re: Qwen3: Think deeper, act faster

#166
post #98

I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.

I similarly have a small, simple spatial reasoning problem that only reasoning models get right, and not all of them, and which Qwen3 on max reasoning still gets wrong. > I put a coin in a cup and slam it upside-down on a glass table. I can't see the coin because the cup is over it. I slide a mirror under the table and see heads. What will I see if I take the cup (and the mirror) away?

Simple Claude 3.5 with no reasoning gets it right.

Re: Qwen3: Think deeper, act faster

#168
post #158

Earlier quoted context omitted.

I similarly have a small, simple spatial reasoning problem that only reasoning models get right, and not all of them, and which Qwen3 on max reasoning still gets wrong. > I put a coin in a cup and slam it upside-down on a glass table. I can't see the coin because the cup is over it. I slide a mirror under the table and see heads. What will I see if I take the cup (and the mirror) away?

My first try (omitting chain of thought for brevity): When you remove the cup and the mirror, you will see tails. Here's the breakdown: Setup: The coin is inside an upside-down cup on a glass table. The cup blocks direct view of the coin from above and below (assuming the cup's base is opaque). Mirror Observation: A mirror is slid under the glass table, reflecting the underside of the coin (the side touching the tabl…

Huh, for me it said:

Answer: You will see the same side of the coin that you saw in the mirror — heads .

Why?

The glass table is transparent , so when you look at the coin from below (using a mirror), you're seeing the top side of the coin (the side currently facing up). Mirrors reverse front-to-back , not left-to-right. So the image is flipped in depth, but the orientation of the coin (heads or tails) remains clear. Since the coin hasn't moved during this process, removing the cup and mirror will reveal the exact same face of the coin that was visible via the mirror — which was heads.

Final Answer: You will see heads.

Re: Qwen3: Think deeper, act faster

#169
The chat is the most annoying page ever. If I must be logged in to test it, then the modal to log in should not have the option "stay logged out". And if I choose the option "stay logged out" it should let me enter my test questions without popping up again and again.

Re: Qwen3: Think deeper, act faster

#170

Earlier quoted context omitted.

Except, very literally, data is a collection of single points (ie what we call "anecdotes").

No, Wittgenstein's rule following paradox, Shannon sampling theorem, the law that infinite polynomials pass through any finite set of points (does that have a name?), etc, etc. are all equivalent at the limit to the idea that no amount of anecdotes-per-se add up to anything other than coincidence

No, no, no. Each of them gives you information.
Post reply on HN