Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

21–30 of 450 posts

Re: Qwen3-Max-Thinking

#21
post #9

Earlier quoted context omitted.

afaiu not all of their models are open weight releases, this one so far is not open weight (?)

What would a good coding model to run on an M3 Pro (18GB) to get Codex like workflow and quality? Essentially, I am running out quick when using Codex-High on VSCode on the $20 ChatGPT plan and looking for cheaper / free alternatives (even if a little slower, but same quality). Any pointers?

Z.ai has glm-4.7. Its almost as good for about $8/mo.

Re: Qwen3-Max-Thinking

#22
post #7

I just wanted to check whether there is any information about the pricing. Is it the same as Qwen Max? Also, I noticed on the pricing page of Alibaba Cloud that the models are significantly cheaper within mainland China. Does anyone know why? https://www.alibabacloud.com/help/en/model-studio/models?spm...

I guess they want to partially subsidize local developers?

Maybe that's a requirement from whoever funds them, probably public money.

Re: Qwen3-Max-Thinking

#23

Mandatory pelican on bicycle: https://www.svgviewer.dev/s/U6nJNr1Z

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

Because no one cares about optimizing for this because it's a stupid benchmark.

It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files.

I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling.

The Pelican on a bicycle thing is funny, yes, but it doesn't really translate into more revenue for AI labs so there's a reason it's not radically improving over time.

Re: Qwen3-Max-Thinking

#25

Mandatory pelican on bicycle: https://www.svgviewer.dev/s/U6nJNr1Z

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

It would be trivial to detect such gaming, tho. That's the beauty of the test, and that's why they're probably not doing it. If a model draws "perfect" (whatever that means) pelicans on a bike, you start testing for owls riding a lawnmower, or crows riding a unicycle, or x _verb_ on y ...

Re: Qwen3-Max-Thinking

#26

Earlier quoted context omitted.

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…

+1 to "it's a stupid benchmark".

Re: Qwen3-Max-Thinking

#27

Aghhh, I wished they release a model which outperforms Opus 4.5 in agentic coding in my earlier comments, seems I should wait more. But I am hopeful

Check out the GLM models, they are excellent

Minimax m2.1 rivals GLM 4.7 and fits in 128GB with 100k context at 3bit quantization.

Re: Qwen3-Max-Thinking

#28
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

This is what I find hilarious when these articles assess "factual" knowledge..

We are at the realm of semantic / symbolic where even the release article needs some meta discussion.

It's quite the litmus test of LLMs. LLMs just carry humanities flaws

Re: Qwen3-Max-Thinking

#30
post #28
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

This is what I find hilarious when these articles assess "factual" knowledge.. We are at the realm of semantic / symbolic where even the release article needs some meta discussion. It's quite the litmus test of LLMs. LLMs just carry humanities flaws

(Edited, sorry.)

Yes, of course LLMs are shaped by their creators. Qwen is made by Alibaba Group. They are essentially one with the CCP.

Post reply on HN