Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

31–40 of 450 posts

Re: Qwen3-Max-Thinking

#31

Earlier quoted context omitted.

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

It would be trivial to detect such gaming, tho. That's the beauty of the test, and that's why they're probably not doing it. If a model draws "perfect" (whatever that means) pelicans on a bike, you start testing for owls riding a lawnmower, or crows riding a unicycle, or x _verb_ on y ...

It could still be special-case RLHF trained, just not up to perfection.

Re: Qwen3-Max-Thinking

#33

Earlier quoted context omitted.

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…

It shows that these are nowhere near anything resembling human intelligence. You wouldn't have to optimize for anything if it would be a general intelligence of sorts.

Re: Qwen3-Max-Thinking

#34
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

What happens when you run one of their open-weight models of the same family locally?

Re: Qwen3-Max-Thinking

#35
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

ask who was responsible for the insurrection on january 6th

Re: Qwen3-Max-Thinking

#36

Earlier quoted context omitted.

Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…

It shows that these are nowhere near anything resembling human intelligence. You wouldn't have to optimize for anything if it would be a general intelligence of sorts.

Here's a pencil and paper. Let's see your SVG pelican.

Re: Qwen3-Max-Thinking

#37
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

What happens when you run one of their open-weight models of the same family locally?

Last time I tried something like that with an offline Qwen model I received a non-answer, no matter how hard I prompted it.

Re: Qwen3-Max-Thinking

#38
post #35
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

ask who was responsible for the insurrection on january 6th

You do it, my IP is now flagged (tried incognito and clearing cookies) - they want to have my phone number to let me continue using it after that one prompt.

Re: Qwen3-Max-Thinking

#39

Earlier quoted context omitted.

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…

Why stupid? Vector images are widely used and extremely useful directly and to render raster images at different scales. It’s also highly connected with spacial and geometric reasoning and precision, which would open up a whole new class of problems these models could tackle. Sure, it’s secondary to raster image analysis and generation, but curious why it would be stupid to persue?

Re: Qwen3-Max-Thinking

#40
I'm not familiar with these open-source models. My bias is that they're heavily benchmaxxing and not really helpful in practice. Can someone with a lot of experience using these, as well as Claude Opus 4.5 or Codex 5.2 models, confirm whether they're actually on the same level? Or are they not that useful in practice?

P.S. I realize Qwen3-Max-Thinking isn't actually an open-weight model (only accessible via API), but I'm still curious how it compares.

Post reply on HN