Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

111–120 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#111
post #88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

What’s the blast radius of this bubble popping? It’s all private investment still right?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#112
post #88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

Meanwhile I can do all that and more with reasonix harness for Deepseek with a cache hit rate of 99%. And that's with unsubsidized American providers like cloudflare or Digital Ocean

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#114
post #102

Earlier quoted context omitted.

VCs are footing the bill for that $200 subscription.

They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)

How can you sanely project the last 12 months forward? We have seen a huge uptick in usage. Last summer AI was a toy to most devs, now every enterprise developer I talked to uses it every day. Coding agent providers are surely going to hit market saturation in the near future.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#115
post #88

Earlier quoted context omitted.

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

What’s the blast radius of this bubble popping? It’s all private investment still right?

Two thirds of most of the DC builds are not compute. So it's a CRE play the last leg holding up that mess.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#116

Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol Source: http://pellmell.ai/leaderboard . This jumps around a lot based on the top throughput and latency of whatever provider happens to be b…

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did.

Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in plain terms.

For a fairly gnarly task, after fighting with with Claude-Code + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it was like a breath of fresh air.

Arguably a key aspect of intelligence is concise, clear communication, and current benchmarks miss that, at least as far as I'm aware. I would think some arena-type benchmarks where humans rate responses would measure this, though I'm not sure which those are.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#117

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

> wayyy overpriced

Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#118

Earlier quoted context omitted.

I have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.

> however, Opus does seem plain fucking stupid Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.

I bought my first LLM subscription with Claude right before they gave access to Fable 5.

I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.

I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#119
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Yeah, I've dropped back to 4.8 entirely for the remainder of this billing cycle. I'm going to be seriously looking into Qwen adoption and harness migration options over the course of August.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#120
post #108
post #88

Earlier quoted context omitted.

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

People keep saying this but from what we’ve seen, Anthropic models are marginally profitable and earn back their costs over their lifetime. The company is burning money building the next versions and other ventures (e.g. verticals), but the models themselves have been profitable.

They're EBITDA profitable, not GAAP profitable.
Post reply on HN