Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

351–360 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#351

Earlier quoted context omitted.

You gotta admit the timing looks very suspicious.

> You gotta admit the timing looks very suspicious. Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model" ? That's indeed a bit fishy.

They could just have avoided all of this by not publishing the benchmark until the new methodology update.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#352

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

And I had the opposite experience. It's a really interesting phenomenon that I can't really explain. My co-founder swears by Deepseek and yet just the other day we were conversing and he was telling me about some of the issues with the way the AI was behaving and trying to show off the cool workarounds he came up with to limit it. I was like, "Interesting, yeah, I've literally never had that problem." I suspect that…

Yeah DeepSeek has always been terrible until the new release version of 4.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#353

Earlier quoted context omitted.

Effective jargon usage is understood by the target audience. If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.

Or you're below its level.

As long as the linguistics are the issue rather than raw inability to understand the concepts being communicated, then the model is just being a bad communicator.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#354

Earlier quoted context omitted.

recommend opencode w/qwen 35B or 27B with MTP. My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.

Any experience with Sleev as replacement of DCP? See https://news.ycombinator.com/item?id=48883538 25 days ago > The Sleev (the project has been renamed to make a startup) creator was shilling their project in the OpenCode Discord. That person is very convinced they have something that no one has ever built before. They focused on token reduction without any real evals for capability impacts. I'm generally against th…

No. the opencode plugin dynamic context pruning is essentially just labeling the context and then summarizing it; you can manually expand the context if you missed something, but most of the time, I'm using it to extend context into different scopes rather than trying to achieve the same goal.

There's a few times it gets too dumb to cope, but in comparison with other coding agents, it seems smarter.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#355
post #347

Earlier quoted context omitted.

Why isn’t it more like building a hardware store using lumber you purchased from a competing hardware store? Or founding a school using an education you obtained at a different school?

Because the companies did it underhandedly without prior consent? Imagine you walked into a hardware store to grab some lumber, didn't pay for it, and the store had to call the cops to swing around your house? That's hardly the typical shopping experience, now is it?

I think that’s my point: none of those things happened, so on what basis can we claim that it’s like they did. Who has to agree with what you’re doing in order for it not to be considered underhanded?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#356

Earlier quoted context omitted.

> however, Opus does seem plain fucking stupid Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.

Fable has spoiled us all.

Whether this is true or not has been keeping me up at night the past week. Daily driving Fable is legit superpowers. Running out of usage and trying to work with literally any other model and everything breaks down because they can’t keep up without constantly tripping and derailing everything, meaning I’m working full time to babysit every judgement call they make instead of flying like a rocket.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#357
post #251

Earlier quoted context omitted.

Or are humans more of a bottleneck than before, because to improve on the most complex problems that demonstrate intelligence you need some way to verify that they are correct. If it's hard for humans to even know if something is correct, wouldn't that slow everything down and simply put limits on the scaling speed of models based on human verification? So instead of relying heavily on human bottlenecks, you focus on…

Very interesting point you make! Before LLMs I had a theory that we cannot make something more intelligent/complex than us. LLMs are certainly more knowledgable, but maybe not more intelligent, arguably. It's possible we're approacing a ceiling indeed. Model capability might be on an asymptote appraching but never quite reaching parity with human intelligence.

There are many things that can have some decent level of automatic verification and those were some of the first for LLMs to excel at, like math and programming. Now, physics involves math, but verification would still often require some form of measurement to make sure that the math relates to the real world meaningfully.

Many extended kinds of verification can be done by LLMs, but they need to be able to follow instructions reliably and agentic task orchestration may be critical to that verification process.

There is no doubt they will surpass us as there is a lot of easy to reason about information that they can verify as incrementally proven by other knowledge. The trick is knowing what can be proven with existing knowledge and what needs human evaluation.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#358
post #356

Earlier quoted context omitted.

Fable has spoiled us all.

Whether this is true or not has been keeping me up at night the past week. Daily driving Fable is legit superpowers. Running out of usage and trying to work with literally any other model and everything breaks down because they can’t keep up without constantly tripping and derailing everything, meaning I’m working full time to babysit every judgement call they make instead of flying like a rocket.

Yeah, tell me about it. Anything else feels like what going to a local coding model used to feel like. Granted it still fucked up on occasion, but like maybe twice a week, not literally every other turn.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#359

Earlier quoted context omitted.

They've been continuously programmed with insane beliefs about China, which is less shocking when you understand what insane beliefs that they've had programmed into them about their neighbors. The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time.

> The world and your neighborhood are full of evil communists and Nazis who are trying to kill you all the time. Don't know about that but your neighbors in Iran in early january happened to be "nice people" who just followed the orders to slaughter 30 000 unarmed civilians. We could talk about the, what 600 000 deaths, including many civilians, in the Ukraine/Russia war. Or we could talk about the number of nice pal…

[dead]

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#360

Earlier quoted context omitted.

>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.

Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.

That’s the thing, the difference is so small that you won’t be wasting an hour per day with a model that’s 95% as good. In fact, you’d notice zero difference most of the days and when you do, maybe it’s an extra 30 minutes.

And the price difference is far greater than $2k/month once the API cost is no longer subsidized.

Would your employer be paying an extra $20k/month to Anthropic if it can save you 2 hours a month?

Post reply on HN