Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

281–290 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#281
post #256
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I thin…

You gotta admit the timing looks very suspicious.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#282

Earlier quoted context omitted.

I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.

They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.

We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.

Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.

You may be living in a place that you think shields you from those but I know the ideologies behind these acts.

The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.

Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#283
post #256

Earlier quoted context omitted.

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I thin…

You gotta admit the timing looks very suspicious.

> You gotta admit the timing looks very suspicious.

Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?

That's indeed a bit fishy.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#284

Earlier quoted context omitted.

They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time. Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians. We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war. Or we could talk about the number of nice pal…

[dead]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#285

Earlier quoted context omitted.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.

Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#286
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Who is going to break it to the Americans that China is more than a slight favorite to win an existential battle over which country is better at math?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#287

Earlier quoted context omitted.

They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time. Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians. We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war. Or we could talk about the number of nice pal…

QED

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#288
post #50

Earlier quoted context omitted.

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

The cursing thing blows my mind. "User is upset? Let's make decisions even faster (ie. more wrong) because clearly that's what they want!" It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?

It's training data might have a ton of examples of people hurrying and screwing more after being yelled at

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#289
post #85

Earlier quoted context omitted.

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and ke…

Glad to see this comment as this has generally been my experience as well. I'm really curious to see why it's so infuriating for others. My best guess is I'm using it more conservatively than most other users in this thread.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#290

Earlier quoted context omitted.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.

- load bearing -
Post reply on HN