Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

101–110 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#102

Earlier quoted context omitted.

It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.

VCs are footing the bill for that $200 subscription.

They have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#103
post #84

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs…

[deleted]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#104
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

I have some internal tests I use for areas where one particular solution/paradigm is dominant but worse.

Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use the newer & better solution.

Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.

Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff the newest versions of Opus 4.7, 4.8 and 5 are worse than useless and become a major dev timesink.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#105

Earlier quoted context omitted.

It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.

VCs are footing the bill for that $200 subscription.

At last, a valid usecase for VCs.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#106
post #88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

I thought they are making a profit on API pricing? A quick Google shows somewhere between 50-70% margins on API inference.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#107
post #84

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs…

Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix.

That vibe coding they brag about as if it was a good thing, it shows.

Take their notation for describing permissions. The docs are not comprehensive, and in practice it doesn't quite work how they describe it.

Or their management of sub-agents. I once lost a sub-agent, it finished and disappeared from UI. Apparently, you can't bring it back yourself: you have to ask the parent agent to do it for you. But the parent was Fable, and I ran out of credits, so I was locked out of using my opus sub-agent because of it.

Or an even more grotesque example: when you paste your claude API token to authorize, it covers characters with *. But it seems like an LLM has hallucinated a limit of API key length and the tail of your key stays visible.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#108
post #88

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

And at that cost they're still not profitable. It's going to be a bumpy road ahead...

People keep saying this but from what we’ve seen, Anthropic models are marginally profitable and earn back their costs over their lifetime. The company is burning money building the next versions and other ventures (e.g. verticals), but the models themselves have been profitable.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#109
post #85
post #41

Earlier quoted context omitted.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.

Statements like this typically come from working on the same setup and context using different models. I actually have that very experience now; I work on something security-adjacent so Fable often drops out, at which point Opus behaves like its lobotomized half-sibling. Pardon me the language, but I can't find a better example to be honest.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#110
post #42

Earlier quoted context omitted.

> It's not enough that it's better? It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.

What status quo? Just look at Openrouter's rankings: https://openrouter.ai/rankings Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.

That's got a significant selection bias. Claude and ChatGPT and Gemini and other subs do not go through openrouter.
Post reply on HN