Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

271–280 of 450 posts

Re: Qwen3-Max-Thinking

#271
post #262

Earlier quoted context omitted.

This looks like it's coming from a separate "safety mechanism". Remains to be seen how much censorship is baked into the weights. The earlier Qwen models freely talk about Tiananmen square when not served from China. E.g. Qwen3 235B A22B Instruct 2507 gives an extensive reply starting with: "The famous photograph you're referring to is commonly known as "Tank Man" or "The Tank Man of Tiananmen Square", an iconic imag…

Difficult to blame them, considering censorship exists in the West too.

What prompt should I run to detect western censorship from a LLM?

Re: Qwen3-Max-Thinking

#272
Can't wait for the benchmark at artificial analysis. Qwen team doesn't seem to have updated the information about this new model yet https://chat.qwen.ai/settings/model. I tried getting an api key from alibabacloud, but the amount of steps from creating an account made me stop, it was too much. It should be this difficult.

Incredible work anyways!

Re: Qwen3-Max-Thinking

#273
post #262

Earlier quoted context omitted.

This looks like it's coming from a separate "safety mechanism". Remains to be seen how much censorship is baked into the weights. The earlier Qwen models freely talk about Tiananmen square when not served from China. E.g. Qwen3 235B A22B Instruct 2507 gives an extensive reply starting with: "The famous photograph you're referring to is commonly known as "Tank Man" or "The Tank Man of Tiananmen Square", an iconic imag…

Difficult to blame them, considering censorship exists in the West too.

nowhere near to China.

In US almost anything could be discussed - usually only unlawful things are censored by government.

Private entities might have their own policies, but government censorship is fairly small.

Re: Qwen3-Max-Thinking

#275
post #262

Earlier quoted context omitted.

Difficult to blame them, considering censorship exists in the West too.

nowhere near to China. In US almost anything could be discussed - usually only unlawful things are censored by government. Private entities might have their own policies, but government censorship is fairly small.

In the US, yes, by the law, in principle.

In practice, you will have loss of clients, of investors, of opportunities (banned from Play Store, etc).

In Europe, on top of that, you will get fines, loss of freedom, etc.

Re: Qwen3-Max-Thinking

#276

One thing I’m becoming curious about with these models are the token counts to achieve these results - things like “better reasoning” and “more tool usage” aren’t “model improvements” in what I think would be understood as the colloquial sense, they’re techniques for using the model more to better steer the model, and are closer to “spend more to get more” than “get more for less.” They’re still valuable, but they op…

i'm no expert, and i actually asked google gemini a similar question yesterday - "how much more energy is consumed by running every query through Gemini AI versus traditional search?" turns out that the AI result is actually on par, if not more efficient (power wise) than traditional search. I think it said its the equivalent power of watching 5 seconds of TV per search. I also asked perplexity to give a report of th…

I’m… deeply suspicious of Gemini’s ability to make that assessment.

I do broadly agree that smaller, better tuned models are likely to be the future, if only because the economics of the large models seem somewhat suspect right now, and also the ability to run models on cheaper hardware’s likely to expand their usability and the use cases they can profitably address.

Re: Qwen3-Max-Thinking

#277
post #253

Earlier quoted context omitted.

You conversely get the same issue if you have no guardrails. Ie: Grok generating CP makes it completely unusable in a professional setting. I don't think this is a solvable problem.

Curious why you use abbreviations ? "CP", "MAP", etc just for such.

I'm lazy

Re: Qwen3-Max-Thinking

#280
post #157

Earlier quoted context omitted.

I've yet to encounter any censorship with Grok. Despite all the negative news about what people are telling it to do, I've found it very useful in discussing controversial topics. I'll use ChatGPT for other discussions but for highly-charged political topics, for example, Grok is the best for getting all sides of the argument no matter how offensive they might be.

grok is indeed one of the most permitting models https://speechmap.ai/labs/

Surprising to see Mistral on top there. I’d imagine EU regulations / culture would require them to not be as free speech friendly.
Post reply on HN