Live data from Hacker News

Qwen3.8 Max now ranked as the best overall model by agentic index

artificialanalysis.ai

361–364 of 364 posts

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#361

Earlier quoted context omitted.

Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.

I find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me. One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B…

How do you do the review cycles? Is this some automated feature of your harness? Do you have a generic prompt for this or ask 27B yourself?

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#363

Earlier quoted context omitted.

What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.

question's phrasing made your judgement obvious

I seriously wasnt judging i maybe shouldve asked llm to frame it better cause i knew it might sound that way, thats why i added: im not judging, merely asking...

Re: Qwen3.8 Max now ranked as the best overall model by agentic index

#364
post #209
post #192

Earlier quoted context omitted.

Is there any model that knows how to smooth an overly literary text over? I find Opus and Fable constantly decorate the documentation they write like a damn 19/20th century writer. We're working with IT stuff yet it writes like it's going to win some Pulitzer prize. It's that one thing I don't get why they can't train them to do properly: I have not encountered a model yet that sticks to the current language of the d…

Not sure how to fully fix this but I remember a session last week where I got so fed up mid way though reading a response that I used the following: "give me this again without jargon invented this session at high density and with a couple (maybe more or less) simple useful ascii diagrams underneath each design" The context is that I was discussing an experimental new idea for my video game review analysis product. D…

I've just seen it write this, I'm still laughing/crying:

> Monitor clipping. review_watched produces PathStatus and nothing else. It never feeds solve. Whatever it does to legs cannot reach the search.

(that's after being told twice to not use shorthand jargon nor reference the code directly)

Post reply on HN