Live data from Hacker News

Qwen3: Think deeper, act faster

qwenlm.github.io

221–230 of 412 posts

Re: Qwen3: Think deeper, act faster

#221

Earlier quoted context omitted.

Except, very literally, data is a collection of single points (ie what we call "anecdotes").

No, Wittgenstein's rule following paradox, Shannon sampling theorem, the law that infinite polynomials pass through any finite set of points (does that have a name?), etc, etc. are all equivalent at the limit to the idea that no amount of anecdotes-per-se add up to anything other than coincidence

Without structural assumptions, there is no necessity - only observed regularity. Necessity literally does not exist. You will never find it anywhere.

Hume figured this out quite a while ago and Kant had an interesting response to it. Think the lack of “necessity” is a problem? Try to find “time” or “space” in the data.

Data by itself is useless. It’s interesting to see peoples’ reaction to this.

Re: Qwen3: Think deeper, act faster

#222
post #98

I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.

Current models are quite far away from human-level physical reasoning (paper below). An upcoming version of models trained on world simulation will probably do much better.

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

https://phybench-official.github.io/phybench-demo/

Re: Qwen3: Think deeper, act faster

#225

Earlier quoted context omitted.

No, Wittgenstein's rule following paradox, Shannon sampling theorem, the law that infinite polynomials pass through any finite set of points (does that have a name?), etc, etc. are all equivalent at the limit to the idea that no amount of anecdotes-per-se add up to anything other than coincidence

Without structural assumptions, there is no necessity - only observed regularity. Necessity literally does not exist. You will never find it anywhere. Hume figured this out quite a while ago and Kant had an interesting response to it. Think the lack of “necessity” is a problem? Try to find “time” or “space” in the data. Data by itself is useless. It’s interesting to see peoples’ reaction to this.

@whatnow37373 — Three sentences and you’ve done what a semester with Kritik der reinen Vernunft couldn’t: made the Hume-vs-Kant standoff obvious. The idea that “necessity” is just the exhaust of our structural assumptions (and that data, naked, can’t even locate time or space) finally snapped into focus.

This is exactly the kind of epistemic lens-polishing that keeps me reloading HN.

Re: Qwen3: Think deeper, act faster

#226
post #38

Something that interests me about the Qwen and DeepSeek models is that they have presumably been trained to fit the worldview enforced by the CCP, for things like avoiding talking about Tiananmen Square - but we've had access to a range of Qwen/DeepSeek models for well over a year at this point and to my knowledge this assumed bias hasn't actually resulted in any documented problems from people using the models. Asid…

What I wonder about is whether these models have some secret triggers for particular malicious behaviors, or if that's possible. Like if you provide a code base that had some hints that the code involves military or government networks, whether the model would try to sneak in malicious but obsfucated code with it's output

Re: Qwen3: Think deeper, act faster

#227
post #93

Earlier quoted context omitted.

Alibaba, I have a huge favor to ask if you're listening. You guys very obviously care about the community. We need an answer to gpt-image-1. Can you please pair Qwen with Wan? That would literally change the art world forever. gpt-image-1 is an almost wholesale replacement of ComfyUI and SD/Flux ControlNets. I can't underscore how big of a deal it is. As such, OpenAI has leapt ahead and threatens to start capturing m…

> That would literally change the art world forever. In what world? Some small percentage up or who knows, and _that_ revolutionized art? Not a few years ago, but now, this. Wow.

Forever, as in for a few weeks… ;-)

Re: Qwen3: Think deeper, act faster

#228
post #168
post #158

Earlier quoted context omitted.

My first try (omitting chain of thought for brevity): When you remove the cup and the mirror, you will see tails. Here's the breakdown: Setup: The coin is inside an upside-down cup on a glass table. The cup blocks direct view of the coin from above and below (assuming the cup's base is opaque). Mirror Observation: A mirror is slid under the glass table, reflecting the underside of the coin (the side touching the tabl…

Huh, for me it said: Answer: You will see the same side of the coin that you saw in the mirror — heads . Why? The glass table is transparent , so when you look at the coin from below (using a mirror), you're seeing the top side of the coin (the side currently facing up). Mirrors reverse front-to-back , not left-to-right. So the image is flipped in depth, but the orientation of the coin (heads or tails) remains clear.…

The question doesn't define which side you're going to look from at the end, so either looking down or up is valid.

Re: Qwen3: Think deeper, act faster

#229
post #152

Excellent release by the Qwen team as always. Pretty much the best open-weights model line so far. In my early tests however, several of the advertised languages are not really well supported and the model is outputting something that only barely resembles them. Probably a dataset quality issue for low-resource languages that they cannot personally check for, despite the “119 languages and dialects” claim.

Indeed, I tried several low-resource Romance languages they claim to support and performance is abysmal.

Which languages?

Re: Qwen3: Think deeper, act faster

#230
post #24

I’m most excited about Qwen-30B-A3B. Seems like a good choice for offline/local-only coding assistants. Until now I found that open weight models were either not as good as their proprietary counterparts or too slow to run locally. This looks like a good balance.

It would be interesting to try, but for the Aider benchmark, the dense 32B model scores 50.2 and the 30B-A3B doesn't publish the Aider benchmark, so it may be poor.
Post reply on HN