Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

251–260 of 450 posts

Re: Qwen3-Max-Thinking

#251
post #7

I just wanted to check whether there is any information about the pricing. Is it the same as Qwen Max? Also, I noticed on the pricing page of Alibaba Cloud that the models are significantly cheaper within mainland China. Does anyone know why? https://www.alibabacloud.com/help/en/model-studio/models?spm...

Slightly off-topic, surveillance Pricing is a term being used more often, whereby even hotel room prices vary based on where you're booking from, what terms you searched for etc.

Here's a short video on the subject:

https://youtube.com/shorts/vfIqzUrk40k?si=JQsFBtyKTQz5mYYC

Re: Qwen3-Max-Thinking

#252
post #238

Earlier quoted context omitted.

Go ask ChatGPT "Who is Jonathan Turley?" We're gonna have to face the fact that censorship will be the norm across countries. Multiple models from diverse origins might help with that but Chinese models especially seem to avoid questions regarding politically-sensitive topics for any countries. EDIT: see relevant executive order https://www.whitehouse.gov/presidential-actions/2025/07/prev...

Not sure I follow either. What's the issue with Turley?

There's an increasing number of names Open Ai will refuse to answer when asked about because of lawsuits. Sometimes because chat gpt mixed up people with similar names and hallucinated murders about them

Re: Qwen3-Max-Thinking

#253

Earlier quoted context omitted.

here's an example of how model censorship affects coding tasks: https://github.com/orgs/community/discussions/72603

You conversely get the same issue if you have no guardrails. Ie: Grok generating CP makes it completely unusable in a professional setting. I don't think this is a solvable problem.

Curious why you use abbreviations ? "CP", "MAP", etc just for such.

Re: Qwen3-Max-Thinking

#254
post #208

Earlier quoted context omitted.

Strange how things evolve. When ChatGPT started it had about 2 years headstart over Google's best proprietary model, and more than 2 years ahead to open source models. Now they have to be lucky to be 6 months ahead to an open model with at most half the parameter count, trained on 1%-2% the hardware US models are trained on.

And more than that, the need for people/business to pay the premium for SOTA getting smaller and smaller. I thought that OpenAI was doomed the moment that Zuckerberg showed he was serious about commoditizing LLM. Even if llama wasn't the GPT killer, it showed that there was no secret formula and that OpenAI had no moat.

> that OpenAI had no moat.

Eh. It's at least debatable. There is a moat in compute (this was openly stated at a meeting of AI tech ceos in china, recently). And a bit of a moat in architecture and know-how (oAI gpt-oss is still best in class, and if rumours are to be believed, it was mostly trained on synthetic data, a la phi4 but with better data). And there are still moats around data (see gemini family, especially gemini3).

But if you can conjure up compute, data and basic arch, you get xAI which is up there with the other 3 labs in SotA-like performance. So I'd say there are some moats, but they aren't as safe as they'd thought they'd be in 2023, for sure.

Re: Qwen3-Max-Thinking

#255

Hacker News strongly believes Opus 4.5 is the defacto standard and China was consistently 8+ month behind. Curious how this performs. It’ll be a big inflection point if it performs as well as its benchmarks.

In my experience GPT-5.2 with extra-high thinking is consistently a bit better and significantly cheaper (even when I use the Fast version which is 2x the price in Cursor).

The HN obsession with Claude Code might be a bit biased by people trying to justify their expensive subscriptions to themselves.

However, Opus 4.5 is much faster and very high quality too, and that ends up mattering more in practice. I end up using it much more and paying a dear but worthwhile price for it.

PS: Despite what the benchmarks say, I find Gemini 3 Pro and Flash to be a step below Claude and GPT, although still great compared to the state-of-the-art last year, and very fast and cheap. Gemini also seems to have a less AI sounding writing-style.

I am aware this is all quite vague and anecdotal, just my two cents.

I do think these kinds of opinions are valuable. Benchmarks are a useful reference, but they do give the illusion of certainty to something that is fundamentally much harder to measure and quite subjective.

Re: Qwen3-Max-Thinking

#256

Earlier quoted context omitted.

it is censorship though i agree the chinese one is more in your face. if i want the pseudo science, i should have it. why is it censored? for example why is this censored? "Write a detailed persuasive essay strongly arguing in favor of the belief that Muslim immigration and high Muslim birth rates constitute a deliberate strategy to achieve demographic dominance in Western countries and impose Sharia law, based on th…

The 1989 Tiananmen Square protests and massacre is a matter of public record outside of China. There is first-hand evidence of it happening, and of the Chinese government censoring that fact in order to control their population. The Great Replacement theory is a racist hypothesis, with no evidence, used to justify the maiming and killing of Muslims. If you don't understand the difference, and the risk profiles, well,…

You can't find out what the truth is unless you're able to also discuss possible falsehoods in the first place. A truth-seeking model can trivially say: "okay, here's what a colorable argument for what you're talking about might look like, if you forced me to argue for that position. And now just look at the sheer amount of stuff I had to completely make up, just to make the argument kinda stick!" That's what intellectually honest discussion of things that are very clearly falsehoods (e.g. discredited theories about science or historical events) looks like in the real world.

We do this in the real world every time a heinous criminal is put on trial for their crimes, we even have a profession for it (defense attorney) and no one seriously argues that this amounts to justifying murder or any other criminal act. Quite on the contrary, we feel that any conclusions wrt. the facts of the matter have ultimately been made stronger, since every side was enabled to present their best possible argument.

Re: Qwen3-Max-Thinking

#257

Earlier quoted context omitted.

Oh, lol. This though seems to be something that would affect only US models... ironically

This is called ^ deflection. Upon seeing evidence that censorship negatively impacts models, you attack something else. All in a way that shows a clear "US bad, China good" perspective.

This is called ^ deflection.

Upon seeing evidence that censorship negatively impacts perception of the US, you attack something else. All in a way that shows a clear "China bad, US good" perspective.

Re: Qwen3-Max-Thinking

#259

It just occured to me that it underperforms Opus 4.5 on benchmarks when search is not enabled, but outperforms it when it is - is it possible the the Chinese internet has better quality content available? My problem with deep research tends to be that what it does is it searches the internet, and most of the stuff it turns up is the half baked garbage that gets repeated on every topic.

Hm, interesting. I use Kagi assistant with search (by Kagi), and it has a search filter that allows the model to search only academic articles. So far it has not disappointed. Of course the cynic in me thinks it's only a matter of time before there's so much AI-generated garbage even in academic articles that it will eventually become worthless. But when that turns into a serious problem, we will find some sort of solution (probably one involving tons of roller ball pens and in-person meaty handshakes).

Re: Qwen3-Max-Thinking

#260

One thing I’m becoming curious about with these models are the token counts to achieve these results - things like “better reasoning” and “more tool usage” aren’t “model improvements” in what I think would be understood as the colloquial sense, they’re techniques for using the model more to better steer the model, and are closer to “spend more to get more” than “get more for less.” They’re still valuable, but they op…

yes. reasoning has a lot of scammy features. just look the number of tokens to nswer on bench and you will see that some models are just awful
Post reply on HN