Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

291–300 of 450 posts

Re: Qwen3-Max-Thinking

#292
post #275

Earlier quoted context omitted.

nowhere near to China. In US almost anything could be discussed - usually only unlawful things are censored by government. Private entities might have their own policies, but government censorship is fairly small.

In the US, yes, by the law, in principle. In practice, you will have loss of clients, of investors, of opportunities (banned from Play Store, etc). In Europe, on top of that, you will get fines, loss of freedom, etc.

Others responding to my speech by exercising their own rights to free speech and free association as individuals does not violate my right to free speech. One can make an argument that corporations doing those things (e.g. your Play Store example) is sufficiently different in kind to individuals doing it -- and a lot of people would even agree with that argument! It does, however, run afoul of current first amendment jurisprudence.

Either way, this is categorically different from China's policies on e.g. Tibet, which is a centrally driven censorship decision whose goal is to suppress factual information.

Re: Qwen3-Max-Thinking

#293
post #262

Earlier quoted context omitted.

This looks like it's coming from a separate "safety mechanism". Remains to be seen how much censorship is baked into the weights. The earlier Qwen models freely talk about Tiananmen square when not served from China. E.g. Qwen3 235B A22B Instruct 2507 gives an extensive reply starting with: "The famous photograph you're referring to is commonly known as "Tank Man" or "The Tank Man of Tiananmen Square", an iconic imag…

Difficult to blame them, considering censorship exists in the West too.

yeah, censorship in the west should give them carte blanche, difficult to blame them, what a fool

Re: Qwen3-Max-Thinking

#294

Earlier quoted context omitted.

This looks like it's coming from a separate "safety mechanism". Remains to be seen how much censorship is baked into the weights. The earlier Qwen models freely talk about Tiananmen square when not served from China. E.g. Qwen3 235B A22B Instruct 2507 gives an extensive reply starting with: "The famous photograph you're referring to is commonly known as "Tank Man" or "The Tank Man of Tiananmen Square", an iconic imag…

I run cpatonn/Qwen3-VL-30B-A3B-Thinking-AWQ-4bit locally. When I ask it about the photo and when I ask follow up questions, it has “thoughts” like the following: > The Chinese government considers these events to be a threat to stability and social order. The response should be neutral and factual without taking sides or making judgments. > I should focus on the general nature of the protests without getting into spe…

Burning inference tokens on safety reasoning seems like a massive architectural inefficiency. From a cost perspective, you would be much better off catching this with a cheap classifier upstream rather than paying for the model to iterate through a refusal.

Re: Qwen3-Max-Thinking

#295

Earlier quoted context omitted.

Why is this surprising? Isn't it mandatory for chinese companies to do adhere to the censorship? Aside from the political aspect of it, which makes it probably a bad knowledge model, how would this affect coding tasks for example? One could argue that Anthropic has similar "censorships" in place (alignment) that prevent their model from doing illegal stuff - where illegal is defined as something not legal (likely?) i…

There's a pretty huge difference between relatively generic stuff like "don't teach people how to make pipe bombs" or whatever vs "don't discuss topics that are politically sensitive specifically in ." The equivalent here for the US would probably be models unwilling to talk about chattel slavery, or Japanese internment, or the Tuskegee Syphilis Study.

The US has plenty of examples of censorship that's politically motivated, particularly around certain medical products.

Re: Qwen3-Max-Thinking

#296

One thing I’m becoming curious about with these models are the token counts to achieve these results - things like “better reasoning” and “more tool usage” aren’t “model improvements” in what I think would be understood as the colloquial sense, they’re techniques for using the model more to better steer the model, and are closer to “spend more to get more” than “get more for less.” They’re still valuable, but they op…

I also find the implications for this for AGI interesting. If very compute-intensive reasoning leads to very powerful AI, the world might remain the same for at least a few years even after the breakthrough because the inference compute simply cannot keep up.

You might want millions of geniuses in a data center, but perhaps you can only afford one and haven't built out enough compute? Might sound ridiculous to the critics of the current data center build-out, but doesn't seem impossible to me.

Re: Qwen3-Max-Thinking

#297
post #283

Earlier quoted context omitted.

> to deny climate change This is a statement of facts, just like the Tiananmen Square example is a statement of fact. What is interesting in the Alibaba Cloud case is that the model output is filtered to remove certain facts. The people claiming some "both sides" equivalence, on the other hand, are trying to get a model to deny certain facts.

“We have facts, they have falsities”. I think the crux of the issue here is that facts don’t exist in reality, they are subjective by their very nature. So we have on one side those who understand this, and absolutists like yourself who believe facts are somehow unimpugnable and not subjective. Well, China has their own facts, you have yours, I have mine, and we can only arrive at a fact by curating experiential even…

If a man beats his wife, and stops her from talking about it, has a man really beaten his wife?

Re: Qwen3-Max-Thinking

#298

> By scaling up model parameters and leveraging substantial computational resources So, how large is that new model?

While Qwen2.5 was pre-trained on 18 trillion tokens, Qwen3 uses nearly twice that amount, with approximately 36 trillion tokens covering 119 languages and dialects. https://qwen.ai/blog?id=qwen3

Thanks for the info, but I don't think it answers the question. I mean, you could train a 20-node network on 36 trillion tokens. Wouldn't make much sense, but you could. So I was asking more about the number of nodes / parameters or GB of file size.

In addition, there seem to be many different versions of Qwen3. E.g. here the list from ollama library: https://ollama.com/library/qwen3/tags

Re: Qwen3-Max-Thinking

#299
post #262

Earlier quoted context omitted.

This looks like it's coming from a separate "safety mechanism". Remains to be seen how much censorship is baked into the weights. The earlier Qwen models freely talk about Tiananmen square when not served from China. E.g. Qwen3 235B A22B Instruct 2507 gives an extensive reply starting with: "The famous photograph you're referring to is commonly known as "Tank Man" or "The Tank Man of Tiananmen Square", an iconic imag…

Difficult to blame them, considering censorship exists in the West too.

Hard to agree. Not even being to say something because it's either illegal or there are systems to erase it instantly, is very different from people dislike (even too radically) you to say something.

Re: Qwen3-Max-Thinking

#300
post #262

Earlier quoted context omitted.

Difficult to blame them, considering censorship exists in the West too.

nowhere near to China. In US almost anything could be discussed - usually only unlawful things are censored by government. Private entities might have their own policies, but government censorship is fairly small.

Oh yes it is. Anything sexual is heavily censored in the west. In particular the US.
Post reply on HN