Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

71–80 of 450 posts

Re: Qwen3-Max-Thinking

#71

Earlier quoted context omitted.

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…

I suspect there is actually quite a bit of money on the table here. For those of us running print-on-demand workflows, the current raster-to-vector pipeline is incredibly brittle and expensive to maintain. Reliable native SVG generation would solve a massive architectural headache for physical product creation.

Re: Qwen3-Max-Thinking

#72
post #69
post #67

Earlier quoted context omitted.

Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.

Are you saying we cannot talk about the bad things the US has done?

No I'm saying we can, unlike how it is in China. Besides that point, I think GP is arguing that China is villinized more than the US.

Re: Qwen3-Max-Thinking

#73

[flagged]

Why is this surprising? Isn't it mandatory for chinese companies to do adhere to the censorship? Aside from the political aspect of it, which makes it probably a bad knowledge model, how would this affect coding tasks for example? One could argue that Anthropic has similar "censorships" in place (alignment) that prevent their model from doing illegal stuff - where illegal is defined as something not legal (likely?) i…

here's an example of how model censorship affects coding tasks: https://github.com/orgs/community/discussions/72603

Re: Qwen3-Max-Thinking

#74
post #38
post #35

Earlier quoted context omitted.

ask who was responsible for the insurrection on january 6th

You do it, my IP is now flagged (tried incognito and clearing cookies) - they want to have my phone number to let me continue using it after that one prompt.

thats even funnier. thanks for the update.

Re: Qwen3-Max-Thinking

#75

Earlier quoted context omitted.

It shows that these are nowhere near anything resembling human intelligence. You wouldn't have to optimize for anything if it would be a general intelligence of sorts.

Here's a pencil and paper. Let's see your SVG pelican.

So you think if would give a pencil and a paper to the model would it do better?

I don't think SVG is the problem. It just shows that models are fragile (nothing new) so even if they can (probably) make a good PNG with a pelican on a bike, and they can make (probably) make some good SVG, they do not "transfer" things because they do not "understand them".

I do expect models to fail randomly in tasks that are not "average and common" so for me personally the benchmark is not very useful (and that does not mean they can't work, just that I would not bet on it). If there are people that think "if an LLM outputted an SVG for my request it means it can output an SVG for every image", there might be some value.

Re: Qwen3-Max-Thinking

#76

[flagged]

Man, the Chinese government must be a bunch of saints that you must go back 35 years to dig up something heinous that they did.

This suggests that the Chinese government recognises that its legitimacy is conditional and potentially unstable. Consequently, the state treats uncontrolled public discourse as a direct threat. By contrast, countries such as the United States can tolerate the public exposure of war crimes, illegal actions or state violence, since such revelations rarely result in any significant consequences. While public outrage may influence narratives or elections to some extent, it does not fundamentally endanger the continuity of power.

I am not sure if one approach is necessarily worse than the other.

Re: Qwen3-Max-Thinking

#78

Earlier quoted context omitted.

Why is this surprising? Isn't it mandatory for chinese companies to do adhere to the censorship? Aside from the political aspect of it, which makes it probably a bad knowledge model, how would this affect coding tasks for example? One could argue that Anthropic has similar "censorships" in place (alignment) that prevent their model from doing illegal stuff - where illegal is defined as something not legal (likely?) i…

here's an example of how model censorship affects coding tasks: https://github.com/orgs/community/discussions/72603

Oh, lol. This though seems to be something that would affect only US models... ironically

Re: Qwen3-Max-Thinking

#79
post #42

Aghhh, I wished they release a model which outperforms Opus 4.5 in agentic coding in my earlier comments, seems I should wait more. But I am hopeful

One of the ways the chinese companies are keeping up is by training the models on the outputs of the American fronteir models. I'm not saying they don't innovate in other ways, but this is part of how they caught up quickly. However, it pretty much means they are always going to lag.

Does the model collapse proof still hold water these days?

Re: Qwen3-Max-Thinking

#80
post #67

Earlier quoted context omitted.

Are you actually defending the censorship of Tiananmen Square?

Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.

The US govt doesn't force censorship of its history, good or bad.
Post reply on HN