Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

131–140 of 372 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#132

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

Damn I thought it was my extra instructions, I swear everything it writes is in some shorthand with direct references to variables that literally nobody could figure out unless you literally just wrote that code 5 minutes ago. I had it stop writing comments altogether cause it was always four lines of complete and utter nonsense, and it doesn't even obey that rule half the time. Despite doing an extensive back and fo…

Everything is load-bearing with 3 measured blockers.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#133
post #44

Earlier quoted context omitted.

Sent to solve one task, came back with half of it solved and 2 more problems.

"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"

This is driving me crazy. opus 4.8 did not do this to me not (at least during pre-5.0 timeframe). Feels like the new cycle is one step forward, two steps back.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#134
post #107

Earlier quoted context omitted.

Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix. That vibe coding they brag about as if it was a good thing, it shows. Take their notation for describing permissions. The docs are not comprehensive, and…

so many ridiculous "how the fuck did this get through basic QA?" issues with Claude Code. I can't believe how many critical bugs fall through. My favourite one is the bug where Plan mode can execute destructive commands inadvertently. Then all these get closed with `Closing for now — inactive for too long. Please open a new issue if this is still relevant.`. Awesome.

> I can't believe how many critical bugs fall through.

Almost like CC is 100% vibe coded.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#135
post #4

Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.

Why does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index.

What cost the most in API. Input, Cached Input, or Output. There you have your answer.

Unfortunately, we have moved so much of the actual intelligence of models towards reasoning, what results in some models getting good scores, but this is because they are dumping a insane amount of reasoning tokens at the problem.

So a mid priced model, with heavy reasoning output, cost the same as a expensive model, with medium reasoning output.

Before the GPT Luna price drop of 80%, you actually had the same price if you used Luna High and Sol Low. With the difference that Sol Low was insane fast, and often way better code.

https://deepswe.datacurve.ai/

Do not look at the top score but more what is on the horizontal axis as you go down. Sol Medium is frankly, was the best performance for dollar, until that Luna price drop. I will even argue that despite the higher price, Sol Medium is still way better despite Luna Max being cheaper. Or Opus Low, one of the better values also.

What do you notice? Is that those models all have a high intelligence start point for their low setting. So that means they do not rely as much on output tokens aka thinking.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#136

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Anthropic is the new AWS.

Amazon's first principle is the Customer Obsession. Making customers happy.

Fun bit is that the human psychology rates personal looking fixes better than having no issues at all.

For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment".

Meanwhile, any other (small) cloud. Simple, no weird charges. Even _most_ of network egress is free. But, no reason to call support or feel "extraordinary". Comes out as "meh" against Amazon's "top tier" support model...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#137

Earlier quoted context omitted.

> wayyy overpriced Maybe they consider that hiring a person to do it would have cost at least as much and taken much more time, so paying them is a bargain.

Yeah, but now we can hire the Chinese instead for 1/100th the cost. It's an even better deal. Plus we get to own, keep, run, do whatever with the model. We don't feel trapped. Moreover, it's something we can truly build on top of and own our own destiny. Anthropic and OpenAI are the new Oracle (Oracle pre-AI; Oracle is even worse now). Expensive, feels like dealing with a lawyer, and not at all open. They just became…

> the new Oracle

I think that's exactly what they are going for - enterprise and government customers.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#138
post #107
post #84

Earlier quoted context omitted.

I managed to lose around $300 in credits I had saved for some emergency /fast sessions the following way: switch to Fable. Work on the design. Downgrade to Opus for the build. If any of other parallel Opus session has /fast enabled it seems to enable it for the newly spawned session by default. Before I knew it, the $300 was gone. I think the bug is now solved, but it was rather unpleasant. I dont ever remember bugs…

Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix. That vibe coding they brag about as if it was a good thing, it shows. Take their notation for describing permissions. The docs are not comprehensive, and…

What amazes me is how, for a vibe coded product where all they have to do is use their AI to fix things ... NOTHING EVER GETS FIXED!

I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment.

IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#140
post #50
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Agreed, it's extremely frustrating. It's the only model that actually makes me curse when talking to it, even knowing how counterproductive it is.

I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard.

With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or maybe just thinks for 10 minutes instead, then fixes one thing and breaks four additional ones.

Post reply on HN