Live data from Hacker News

Qwen3-4B-Thinking-2507

huggingface.co

1–10 of 64 posts

Re: Qwen3-4B-Thinking-2507

#4
According to the benchmarks, this one is improved in every one of them compared to the previous version, some better than 30B-A3B. Definitely worth a try, it’ll easily fit into memory and token generation speed will be pleasantly fast.

Re: Qwen3-4B-Thinking-2507

#7
post #6

Is there like a leaderboard or power rankings sort of thing that tracks these small open models and assigns ratings or grades to them based on particular use cases?

https://artificialanalysis.ai/leaderboards/models?open_weigh...

Compare these rankings to actual usage: https://openrouter.ai/rankings

Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance?

Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops.

My interpretation is that the models at the top of ArtificialAnalysis are focusing the most on public benchmarks in their training. Note I am not saying XAI is necessarily nefariously doing this, could just be that they decided it's better bang for the buck to rely on public benchmarks than to try to focus on building their own evaluation systems.

But Grok is not very good compared to the anthropic, openai, or google models despite ranking so highly in benchmarks.

Re: Qwen3-4B-Thinking-2507

#8
post #7
post #6

Earlier quoted context omitted.

https://artificialanalysis.ai/leaderboards/models?open_weigh...

Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…

Claude Opus is in the top 10, also people via OpenRouter mostly use these models for coding and Claude models are particularly good at this, the benchmark doesn't account only for coding capacities tho

Re: Qwen3-4B-Thinking-2507

#9
post #7
post #6

Earlier quoted context omitted.

https://artificialanalysis.ai/leaderboards/models?open_weigh...

Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…

The openrouter rankings can be biased.

For example, Google's inexplicable design decisions around libraries and APIs means it's often worth the 5% premium to just use OpenRouter to access their models. In other cases it's about which models particular agents default to.

Sonnet 4 is extremely good for tool-usage agentic setups though - something I have found other models struggle to do over a long-context.

Re: Qwen3-4B-Thinking-2507

#10

According to the benchmarks, this one is improved in every one of them compared to the previous version, some better than 30B-A3B. Definitely worth a try, it’ll easily fit into memory and token generation speed will be pleasantly fast.

There is a new Qwen3-30B-A3B, you are compare it to the old one. https://huggingface.co/Qwen/Qwen3-30B-A3B-Thinking-2507
Post reply on HN