Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

311–320 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#312

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

If they already require your constant supervision the reason is money.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#313

Earlier quoted context omitted.

As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.

> They didn’t accept guilt for incorporating all of human output into their model without consent. Because that use case is actually permitted by law.

I mean... that's one interpretation of the law, sure.

The law was written before the idea of an LLM existed, and some judges in some specific cases decided the previous law covered this usage.

So, it comes down to if you believe a couple judges ruling on a couple cases is the right way to determine a world-altering new legal framework.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#314

Earlier quoted context omitted.

As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.

Copyright law only considers illegal ownership of a work, so the crime - or tort - was making/acquiring copies without permission or payment. Training from copies has been ruled fair use because it's "transformative" and not simply "derivative." This is obviously debatable, but that's where the debate is at the moment.

> but that's where the debate is at the moment.

Because of the rulings of a couple of judges. Is that actually what the majority of people think?

> Copyright law only considers illegal ownership of a work

That's definitely not true. File sharing, for example, is illegal even if you legally own the original copy you're sharing.

Similarly, copyright has something to say if I read a legal copy of harry potter and then create a new work in that world.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#315

Earlier quoted context omitted.

Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.

They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.

> LLM providers can distill all of human output into their models for 'free'

Not sure what part of being charged guilty and paying a fine you see as "free".

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#316

Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…

I had same experience with OpenAI. I have the $200/month plan and use 5.6 Sol all the time. What would normally use about 2% of my weekly allowance burned through $100 of credits in 40 minutes.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#317
post #34

Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.

It's my daily driver. I like it and find it noticeably better than Opus 4.8. After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote wh…

> My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.

I’m doing the same right now, and I’ve found that asking for “simple English” works most of the times, although not always.

Did you find better wording that works consistently?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#319

Earlier quoted context omitted.

All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…

> and dense with cryptic language Yeah if you ask it a medical question, it answers in that impenetrable jargony style that clinical journals use... full of unnecessarily custom adjectives ("orthopedic" instead of "of the bone") and discipline-specific terms (anterior, distal) even when the user didn't display mastry of this terminology (hint to frontier labs: add training cases for this; it will improve your model's…

it's like asking a developer to explain something.

I always get Haiku to rephrase anything human-facing.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#320
post #256

Earlier quoted context omitted.

Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date. The order changes but I thin…

You gotta admit the timing looks very suspicious.

Luna pricing was just cut by 80% https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog post states is a more accurate judge than the previous methodology.
Post reply on HN