Live data from Hacker News

Qwen3-4B-Thinking-2507

huggingface.co

11–20 of 64 posts

Re: Qwen3-4B-Thinking-2507

#12
post #7
post #6

Earlier quoted context omitted.

https://artificialanalysis.ai/leaderboards/models?open_weigh...

Compare these rankings to actual usage: https://openrouter.ai/rankings Claude is not cheap, why is it far and away the most popular if it's not top 10 in performance? Qwen3 235b ranks highest on these benchmarks among open models, but I have never met someone who prefers its output over Deepseek R1. It's extremely wordy and often gets caught in thought loops. My interpretation is that the models at the top of Artific…

Thanks for sharing that. Interesting that the leaderboard is dominated by Anthropic, Google and DeepSeek. Openai doesn't even register.

Re: Qwen3-4B-Thinking-2507

#14
Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.

Re: Qwen3-4B-Thinking-2507

#15
post #13
post #3

This one should work on personal computers! I'm thankful for Chinese companies raising the floor.

[flagged]

I’m American. Just giving some background to the feeling. There’s some discontent with some western communities (localllama) that Chinese developers have been open weighting all of their models while most western models have been closed weights.

Re: Qwen3-4B-Thinking-2507

#16
post #13
post #3

This one should work on personal computers! I'm thankful for Chinese companies raising the floor.

[flagged]

Let's say any country create the most powerful - and thus best - LLMs. They over time infiltrate it with their political will. Over 20-30 years, I'd imagine people asking those LLMs will have their minds' shifted.

But. That's just me, my pessimism-sci-fi scenario.

Re: Qwen3-4B-Thinking-2507

#17
post #14

Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.

This has been around for a while https://lmarena.ai/leaderboard/text/coding

Re: Qwen3-4B-Thinking-2507

#18
post #14

Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.

openrouter usage stats

Re: Qwen3-4B-Thinking-2507

#20
post #18
post #14

Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.

openrouter usage stats

https://openrouter.ai/rankings

The new qwen3 model is not out yet.

Post reply on HN