Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

231–240 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#231
post #44
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

Sent to solve one task, came back with half of it solved and 2 more problems.

What really enrages me is the amount of effort it puts into justifying weaseling out of work. (THAT'S MY JOB!)

It will do everything it can to defer or push it off, to the point where I’ve had to add multiple imperative directives to the AGENTS file telling it, in no uncertain terms, not to defer tasks under any circumstances.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#232

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

Anthropic is the new AWS. Amazon's first principle is the Customer Obsession. Making customers happy. Fun bit is that the human psychology rates personal looking fixes better than having no issues at all. For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment". Meanwhile,…

Aws is an infrastructure company that builds services on top of that infra to sell more of it at a higher margin.

Anthropic trains models on AWS's (and GCPs, and Microslop's) infrastructure, then skims margin off of selling inference also on the infrastructure owned by the other companies.

These are extremely different businesses.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#233

Earlier quoted context omitted.

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

I gathered that the most recent advances haven’t been in capabilities of the model but more the way that it’s able to be employed (most recently agents).

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#234
post #212

The fact that the Chinese models have caught up on benchmarks suggests to me that its likely we will start to transition now into much more of a brand war. It will be subjective qualities that drive our decisions more than measures of absolute intelligence. Already I am choosing models more because I like the personality or style of what they do than because I think they have the absolute highest chance of outputting…

For me now it’s simply cost and speed. With GPT5.6 and Fable (and respective open models since then) we passed a threshold where intelligence is sufficient. Now I just need speed of iteration and good prices.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#235
post #152

Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?

Apple is already doing this... they worked with Gemini to distill the model into a smaller one that fits on your phone. If you have iOS 27 Beta, you're already using this

sort of. they have a local model, it does some things. they also have significant cloud infrastructure backing it, and most tasks are going to be sent off to the cloud for processing, not be handled by the on-device model. Siri is not on-device by any stretch of the imagination.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#237
post #166

China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing.

When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left

And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#238

Earlier quoted context omitted.

If the legal system declares the first thief’s theft not theft then all bets are off.

> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.

As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#239
post #74
post #71

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Ana…

They JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-i...

> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#240
post #5

I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth. I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A vers…

Did you also try Opus 5 and 5.6 Sol?

Opus 5 is just terrible
Post reply on HN