Live data from Hacker News

Qwen3-Max-Thinking

qwen.ai

61–70 of 450 posts

Re: Qwen3-Max-Thinking

#61
It just occured to me that it underperforms Opus 4.5 on benchmarks when search is not enabled, but outperforms it when it is - is it possible the the Chinese internet has better quality content available?

My problem with deep research tends to be that what it does is it searches the internet, and most of the stuff it turns up is the half baked garbage that gets repeated on every topic.

Re: Qwen3-Max-Thinking

#63
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

Go ahead and ask ChatGPT who Jonathan Turley is, you'll get a similar error "Unable to process response".

It turns out "AI company avoids legal jeopardy" is universal behavior.

Re: Qwen3-Max-Thinking

#64

Mandatory pelican on bicycle: https://www.svgviewer.dev/s/U6nJNr1Z

Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.

It’d be difficult to use in any automated process, as the judgement for how good one of these renditions is, is very qualitative.

You could try to rasterize the SVG and then use an image2text model to describe it, but I suspect it would just “see through” any flaws in the depiction and describe it as “a pelican on a bicycle” anyway.

Re: Qwen3-Max-Thinking

#66
post #24

I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."

Go ahead and ask ChatGPT who Jonathan Turley is, you'll get a similar error "Unable to process response". It turns out "AI company avoids legal jeopardy" is universal behavior.

This one seems to be related to an individual who was incorrectly smeared by chatgpt. (Edited.)

> The AI chatbot fabricated a sexual harassment scandal involving a law professor--and cited a fake Washington Post article as evidence.

https://www.washingtonpost.com/technology/2023/04/05/chatgpt...

That is way different. Let's review:

a) The Chinese Communist Party builds an LLM that refuses to talk about their previous crimes against humanity.

b) Some americans build an LLM. They make some mistakes - their LLM points out an innocent law professor as a criminal. It also invent a fictitious Washington Post article.

The law professor threatens legal action. The american creators of the LLM begin censoring the name of the professor in their service to make the threat go away.

Nice curveball though. Damn.

Re: Qwen3-Max-Thinking

#67

Earlier quoted context omitted.

Man, the Chinese government must be a bunch of saints that you must go back 35 years to dig up something heinous that they did.

Are you actually defending the censorship of Tiananmen Square?

Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.

Re: Qwen3-Max-Thinking

#68
post #42

Aghhh, I wished they release a model which outperforms Opus 4.5 in agentic coding in my earlier comments, seems I should wait more. But I am hopeful

One of the ways the chinese companies are keeping up is by training the models on the outputs of the American fronteir models. I'm not saying they don't innovate in other ways, but this is part of how they caught up quickly. However, it pretty much means they are always going to lag.

They are. There is no way to lead unless China has access to as much compute power.

Re: Qwen3-Max-Thinking

#69
post #67

Earlier quoted context omitted.

Are you actually defending the censorship of Tiananmen Square?

Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.

Are you saying we cannot talk about the bad things the US has done?
Post reply on HN