Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

141–150 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#141

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Its tough to go from max account at home and pay per usage enterprise account at work with heavy usage limits... but the limits are there because pricing is insane. Feel like I'm in the $5 Uber rides phase at home.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#142

Out of curiosity, what's currently the best model I can use locally?

With an unlimited budget, Kimi K3 (which is quite comparable to this Qwen Max imo). With a normal budget/a PC you might already have, probably Qwen 3.6 27B.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#143

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Anthropic is the new AWS. Amazon's first principle is the Customer Obsession. Making customers happy. Fun bit is that the human psychology rates personal looking fixes better than having no issues at all. For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment". Meanwhile,…

Anthropic's constant changing of its mind leads to instability which leads to unhappy customers

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#144
post #110

Earlier quoted context omitted.

What status quo? Just look at Openrouter's rankings: https://openrouter.ai/rankings Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.

That's got a significant selection bias. Claude and ChatGPT and Gemini and other subs do not go through openrouter.

Not really, because that's not a unique aspect of any of those. It's true of all subscription services (that I'm aware of), as well as all of the free models. The selection bias primarily will be against models which be an outlier in the difference between openrouter users and total users, which is a much harder position to argue for any given company except for maybe Twitter.

You can argue there's a selection bias that openrouter users are less likely to display model loyalty, but it would still be a visible confounding factor if it was a statistically significant behavior. And it's not. Nor is there a visibly meaningful indication that people don't sidegrade between models. With every single data set, you're going to see that. You're also going to see it reflected in discourse, as I mentioned. Fact of the matter is there isn't a status quo in AI any more than there's a status quo in cars.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#145

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

You plugged in a space heater on a roofless house.

There is some element of responsibility on the user to guide and monitor the model/harness and not let it rip to burn tokens.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#146
post #141

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Its tough to go from max account at home and pay per usage enterprise account at work with heavy usage limits... but the limits are there because pricing is insane. Feel like I'm in the $5 Uber rides phase at home.

The Chinese models are the public transport in the uber analogy. Once the price the goes up catch the bus!

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#147
post #50
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

The cursing thing blows my mind. "User is upset? Let's make decisions even faster (ie. more wrong) because clearly that's what they want!"

It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#148

Earlier quoted context omitted.

I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.

> however, Opus does seem plain fucking stupid Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.

Fable has spoiled us all.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#149
post #88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

I think profitability is a matter of accounting. Inference is where money is made, but training is where money is spent. We keep getting new models every few months, but frankly the old models are still quite usable. I suspect labs will soon start specializing in expert models per use case so they can increase the lifespan of individual models, and change the profitability per model.
Post reply on HN