Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

361–370 of 371 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#361

Earlier quoted context omitted.

Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.

I find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me. One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B…

How do you do the review cycles? Is this some automated feature of your harness? Do you have a generic prompt for this or ask 27B yourself?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#363

Earlier quoted context omitted.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

question's phrasing made your judgement obvious

I seriously wasnt judging i maybe shouldve asked llm to frame it better cause i knew it might sound that way, thats why i added: im not judging, merely asking...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#364
post #209
post #192

Earlier quoted context omitted.

Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the d…

Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following: "give me this again without jargon invented this session at high density and with a couple (maybe more or less) simple useful ascii diagrams underneath each design" The context is that I was discussing an experimental new idea for my video game review analysis product. D…

I've just seen it write this, I'm still laughing/crying:

> Monitor clipping. review_watched produces PathStatus and nothing else. It never feeds solve. Whatever it does to legs cannot reach the search.

(that's after being told twice to not use shorthand jargon nor reference the code directly)

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#365

Earlier quoted context omitted.

It's well known 35b is much faster (on any hardware) and quite a bit dumber

I had no idea. Where can we learn stuff like this?

Learn about MoE models and also just look at the benchmarks of the two models. For example https://artificialanalysis.ai/models/comparisons/qwen3-6-27b... clearly shows both the intelligence and speed differences. It's a bit degenerate but you can get some useful info from e.g. the /r/localllama subreddit

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#366

Earlier quoted context omitted.

>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.

Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.

Imagine if we had this attitude for electricity. No point building a grid, just give it to a few factories that need all day lighting.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#367

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

I think americans assume when they see a chinese or asian person working at an american business that they "escaped" china as opposed to just being rich enough to go to school abroad. and has little to no bearing on the amount of intelligent going around.

I will just assume they are American, unless they are an uber driver. Some cities in China very much are losers of the new tech plan. Otherwise why would they fly to mexico and cross the border

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#368

Earlier quoted context omitted.

What do you mean by 'Setting the memory to "fast timings"'? The only runtime I can get working for my GPUs is llama.cpp, which I haven't seen anything like that in its argument set. My perusal of the options for vllm and sglang didn't suggest anything similar either before failing miserably.

I think they are referring to the AMD drivers on windows, under the overclocking section you can enable fast memory timings. Not sure if this sort of thing is exposed on linux.

You can edit sys files or use amdmemorytweak on linux.

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#369

Earlier quoted context omitted.

Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?

you seem to be mistaken, Opus and Sol are the worse models.

Stop trying to treat these things like a replacement for yourself and instead approach them as a limitless number of offshore developers. If you are willing to be endearingly literal they will make you happy

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#370
post #346

Earlier quoted context omitted.

Where did that $100B figure come from? I thought they were at ~10B at the end of 2025, so they're either not at 100B yet, or they're growing way faster than 10x / year.

Good question, I heard it on a podcast, but going back to the transcript, looks like that's their forecast, not that they've hit it, they estimated a current $70B, but they've been revising their forecasts up, so yeah, it's probably >10x. Latest solid number they reported was $47B in May.

Wild.
Post reply on HN