=====
LiveCodeBench
E4B IT: 13.2
Qwen: 55.2
===== AIME25
E4B IT: 11.6
Qwen: 81.3
11–20 of 64 posts
=====
LiveCodeBench
E4B IT: 13.2
Qwen: 55.2
===== AIME25
E4B IT: 11.6
Qwen: 81.3
Earlier quoted context omitted.
https://artificialanalysis.ai/leaderboards/models?open_weigh...
Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…
This one should work on personal computers! I'm thankful for Chinese companies raising the floor.
This one should work on personal computers! I'm thankful for Chinese companies raising the floor.
[flagged]
This one should work on personal computers! I'm thankful for Chinese companies raising the floor.
[flagged]
But. That's just me, my pessimism-sci-fi scenario.
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.
So this 4B dense model gets very similar performance to the 30B MoE variant with 7.5x smaller footprint.
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.
openrouter usage stats
The new qwen3 model is not out yet.