Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

191–200 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#191
post #107

Earlier quoted context omitted.

Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix. That vibe coding they brag about as if it was a good thing, it shows. Take their notation for describing permissions. The docs are not comprehensive, and…

What amazes me is how, for a vibe coded product where all they have to do is use their AI to fix things ... NOTHING EVER GETS FIXED! I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment. IF YOU ARE GOING TO HAVE A SHITTY VIBE…

[dead]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#192
post #177

Earlier quoted context omitted.

I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard. With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or ma…

yep matches my experience completely But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.

Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the domain it's tasked with.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#193

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Anthropic is the new AWS. Amazon's first principle is the Customer Obsession. Making customers happy. Fun bit is that the human psychology rates personal looking fixes better than having no issues at all. For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment". Meanwhile,…

I’ve never gotten a refund from Athropic.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#194

Earlier quoted context omitted.

No, a smart model should also give a concise executive summary, "brevity is the soul of wit".

So, in short, once models get smart enough they stop bothering telling us what they did. Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.

Which would be fine, but the parents also just smeared tomato sauce over the walls too, so let’s not get too ahead of ourselves.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#195

Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?

DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0]. [0] https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.

Indeed. I like using Macs mostly, and the bargain M1 Max MBP I am using for local LLMs is a fabulous experimentation platform and does loads of other stuff well, so I am in no rush, but if I reached the point of buying dedicated hardware for an LLM, I'd be looking at the DGX Spark machines.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#196

Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol Source: http://pellmell.ai/leaderboard . This jumps around a lot based on the top throughput and latency of whatever provider happens to be b…

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

Eh I don't know, I care whether it gets the job done and I can see the difference when I review the code, not how well it needs to explain the code to me, I can just read it myself.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#198
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#199
post #44

Earlier quoted context omitted.

Sent to solve one task, came back with half of it solved and 2 more problems.

"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"

And one of them is always something just completely out of scope and the other is something obvious it missed.

“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#200
post #41
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.

> I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff.

To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.

Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".

As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.

Post reply on HN