Earlier quoted context omitted.
I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
Weird how different people's experiences are. If it's making simple mistakes something must be wrong in your setup/context I assume? It's been solid for me, beyond the usual LLMisms that all models have. But I keep context pretty minimal.
Qwen3.8 Max now ranked as the best overall model by agentic index
291–300 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#292Earlier quoted context omitted.
If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.
- load bearing -
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#293Earlier quoted context omitted.
Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?
>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#294Earlier quoted context omitted.
Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#295Earlier quoted context omitted.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
If you talk to the Chinese models, even super smart Qwen 3.8, you can tell they are distilled just from the verbal ticks they have. Gemini, ChatGPT and Claude do not sound alike. The Chinese models 100% sound like one of the 3, usually Claude. American models are load bearing for this LLM generation seam.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#296Earlier quoted context omitted.
Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing. When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple ra…
Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?
Self-hosting is the biggest reason.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#297Earlier quoted context omitted.
Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?
>So why would I want to switch to even worse model? There would be no reason to if you are in the privileged position where cost isn't an issue. For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#298Earlier quoted context omitted.
> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.
Because that use case is actually permitted by law.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#299Earlier quoted context omitted.
> If the legal system declares the first thief’s theft not theft But they didn't find it. The Big LLM provider accepted guilt and paid a fine. You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
> But they didn't find it. The Big LLM provider accepted guilt and paid a fine. That's not how it works. You have to give it back. Otherwise, the distiller can just pay a fine (no larger than the original did) and be okay then, right ?
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#300Earlier quoted context omitted.
"As you requested, I've finished task X. Honestly, task X turned out to require task Y, which I haven't actually done. Task Y is the next step if you'd like to continue along this route."
This is the hard-won load-bearing quote.