I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…
Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I thin…
Qwen3.8 Max now ranked as the best overall model by agentic index
281–290 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#282Earlier quoted context omitted.
I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.
They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.
Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians.
We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war.
Or we could talk about the number of nice palestinians killed since the beginning of the war in Gaza. Or we could go a bit further and talk about the joy and celebration in Gaza after their heroes brought back 200 hostages after having slaughtered 1200 civilians.
You may be living in a place that you think shields you from those but I know the ideologies behind these acts.
The fallacy of gray is just that: it's not true that there's always a nice middle ground and that there's no evil ideology out there.
Something something about the price of liberty being eternal vigilance. For there are people abusing your blind trust.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#283Earlier quoted context omitted.
Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I thin…
You gotta admit the timing looks very suspicious.
Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?
That's indeed a bit fishy.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#284Earlier quoted context omitted.
They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.
> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time. Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians. We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war. Or we could talk about the number of nice pal…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#285Earlier quoted context omitted.
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#286China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#287Earlier quoted context omitted.
They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.
> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time. Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians. We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war. Or we could talk about the number of nice pal…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#288Earlier quoted context omitted.
Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.
The cursing thing blows my mind. "User is upset? Let's make decisions even faster (ie. more wrong) because clearly that's what they want!" It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#289Earlier quoted context omitted.
Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.
Every model that comes out comes with a bunch of people saying "this one is actually dumb they were smart before" and I don't really get it. The models since Opus 4.5 have all been basically the same to me. Sometimes they do the wrong thing, so you have to steer and stop and correct them. Leaving them to operate on their own in no-human-in-the-loop harnesses often gets bad results. But if you single thread it, and ke…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#290Earlier quoted context omitted.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.