China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…
Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
Qwen3.8 Max now ranked as the best overall model by agentic index
221–230 of 364 posts
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#222Earlier quoted context omitted.
It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
VCs are footing the bill for that $200 subscription.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#223Earlier quoted context omitted.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#224Earlier quoted context omitted.
But if you have two experts in a field talking to each other you wouldn't expect them to dumb down their communication.
Effective jargon usage is understood by the target audience. If the AI is communicating to me and can't select the appropriate jargon level, it's failing at communicating effectively.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#225Anthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge…
It’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#226Earlier quoted context omitted.
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
If the legal system declares the first thief’s theft not theft then all bets are off.
But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#227China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 th…
Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#228Earlier quoted context omitted.
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
If the legal system declares the first thief’s theft not theft then all bets are off.
Re: Qwen3.8 Max now ranked as the best overall model by agentic index
#229I believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth. I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A vers…
Did you also try Opus 5 and 5.6 Sol?