Earlier quoted context omitted.
Claude code is just pool quality. They don't make how this thing will behave clear to the user, or give control. They fail at anything that needs an abstraction or model, not just APIs and shell scripts glued together. And "just ask AI" seems to be the default fix. That vibe coding they brag about as if it was a good thing, it shows. Take their notation for describing permissions. The docs are not comprehensive, and…
What amazes me is how, for a vibe coded product where all they have to do is use their AI to fix things ... NOTHING EVER GETS FIXED! I've probably gone to file 20 bugs. In all 20 cases there wasn't just one issue already filed for it: there were several, each which had a bunch of upvotes. And in all 20 cases ... every. last. one. ... Anthropic closed the ticket with no comment. IF YOU ARE GOING TO HAVE A SHITTY VIBE…
Qwen3.8 Max now ranked as the best overall model by agentic index
191–200 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#192Earlier quoted context omitted.
I'd certainly rank it at the very top of the want to kill yourself when using it benchmark. It outperforms everything else on that leaderboard. With weaker models you can sort of understand, they're trying their best and failing, but this thing just channels its immense inteligence into being as annoying as possible instead. I know it can do what I'm asking it to do, but it just finds a way to weasel out of it, or ma…
yep matches my experience completely But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#193Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…
Anthropic is the new AWS. Amazon's first principle is the Customer Obsession. Making customers happy. Fun bit is that the human psychology rates personal looking fixes better than having no issues at all. For example, AWS overcharges you, you contact support, and more or less hassle free they refund or issue credits. The customer feels appreciated, or at least got something "extra" or "special treatment". Meanwhile,…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#194Earlier quoted context omitted.
No, a smart model should also give a concise executive summary, "brevity is the soul of wit".
So, in short, once models get smart enough they stop bothering telling us what they did. Yeah, makes sense. A parent wouldn't bother explaining the details of their job to a toddler.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#195Could someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0]. [0] https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#196Opus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol Source: http://pellmell.ai/leaderboard . This jumps around a lot based on the top throughput and latency of whatever provider happens to be b…
All these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in pl…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#197Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#198China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#199Earlier quoted context omitted.
Sent to solve one task, came back with half of it solved and 2 more problems.
"One thing worth your attention", "Two things worth knowing", "One thing to eyeball"
“One thing worth your attention, if you were to detonate a pipe bomb in your house, it would have a negative effect on your living room”.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#200Any benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
To me it's not so much the dumb mistakes (although there are some of those) but the ultra-verbose, mega-inefficient "solutions" to some problems / prompts.
Stuff that "works" if you're the kind of person that considers slamming a semi-trailer at 200 mph into a door did, technically, result in the door being somehow "open".
As it's supposed to be one of the most advanced model, I can't help but wonder if the solutions are that bad/verbose/inefficient because we're already in a loop of models being trained on sloppy-pasta from previous models.