Live data from Hacker News

Qwen3-4B-Thinking-2507

huggingface.co

41–50 of 64 posts

Re: Qwen3-4B-Thinking-2507

#41
post #40
post #36

If you want to have an opinion on it, just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507... . you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/ Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram. Anyway if you're on mac you can search for "qwen3 4b 2507 mlx…

how about on apple silicon for the iphone

https://joejoe1313.github.io/2025-05-06-chat-qwen3-ios.html

Re: Qwen3-4B-Thinking-2507

#42
post #36

If you want to have an opinion on it, just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507... . you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/ Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram. Anyway if you're on mac you can search for "qwen3 4b 2507 mlx…

Thank you. To spare Mac readers time:

mlx 4bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 5bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 6bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

mlx 8bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

edit: corrected the 4b link

Re: Qwen3-4B-Thinking-2507

#43
post #18

Earlier quoted context omitted.

openrouter usage stats

Since the ranking is based on token usage, wouldn't this ranking be skewed by the fact that small models' APIs are often used for consumer products, especially free ones? Meanwhile reasoning models skew it in the opposite direction, but to what extent I don't know. It's an interesting proxy, but idk how reliable it'd be.

Also, these small models are meant to be run local so not going to appear on openrouter...

Re: Qwen3-4B-Thinking-2507

#44
post #42
post #36

If you want to have an opinion on it, just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507... . you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/ Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram. Anyway if you're on mac you can search for "qwen3 4b 2507 mlx…

Thank you. To spare Mac readers time: mlx 4bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-... mlx 5bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-... mlx 6bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-... mlx 8bit: https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-... edit: corrected the 4b link

Did you mean mlx 4bit:

https://huggingface.co/lmstudio-community/Qwen3-4B-Thinking-...

Re: Qwen3-4B-Thinking-2507

#45
post #36

If you want to have an opinion on it, just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507... . you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/ Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram. Anyway if you're on mac you can search for "qwen3 4b 2507 mlx…

> if you run it at the full 262144 tokens of context youll need ~65gb of ram

What is the relationship between context size and RAM required? Isn't the size of RAM related only to number of parameters and quantization?

Re: Qwen3-4B-Thinking-2507

#46
post #25
post #14

Is there a crowd-sourced sentiment score for models? I know all these scores are juiced like crazy. I stopped taking them at face value months ago. What I want to know is if other folks out there actually use them or if they are unreliable.

Besides the LM Arena Leaderboard mentioned by a sibling comment, if go to the r/LocalLlama/ subreddit, you can very unscientifically get a rough sentiment of the performance of the models by reading the comments (and maybe even check the upvotes). I think the crowd's knee-jerk reaction is unreliable though, but that's what you asked for.

Not anymore tho. It used to be the place to vibe-check a model ~1 year ago, but lately it's filled with toxic my team vs. your team, memes about CEOs (wtf) and general poor takes on a lot of things.

For a while it was china vs. world, but lately it's even more divided, with heavy camping on specific models. You can still get some signal, but you have to either ban a lot of accounts, or read new during different tzs so you can get some of that "i'm just here for the tech stack" vibe from posters.

Re: Qwen3-4B-Thinking-2507

#47
post #6

Is there like a leaderboard or power rankings sort of thing that tracks these small open models and assigns ratings or grades to them based on particular use cases?

https://artificialanalysis.ai/leaderboards/models?open_weigh...

Qwen3-30A-A3B-2507 is much faster on my machine compared to gpt-oss-20B. This leaderboard does not reflect that.

Re: Qwen3-4B-Thinking-2507

#48
post #45
post #36

If you want to have an opinion on it, just install lmstudio and run the q8_0 version of it i.e. here https://huggingface.co/bartowski/Qwen_Qwen3-4B-Instruct-2507... . you can even run it on a 4gb raspberry pi Qwen_Qwen3-4B-Instruct-2507-Q4_K_L.gguf https://lmstudio.ai/ Keep in mind if you run it at the full 262144 tokens of context youll need ~65gb of ram. Anyway if you're on mac you can search for "qwen3 4b 2507 mlx…

> if you run it at the full 262144 tokens of context youll need ~65gb of ram What is the relationship between context size and RAM required? Isn't the size of RAM related only to number of parameters and quantization?

No. Your KV cache is kept in memory also.

Re: Qwen3-4B-Thinking-2507

#49
post #6

Is there like a leaderboard or power rankings sort of thing that tracks these small open models and assigns ratings or grades to them based on particular use cases?

https://artificialanalysis.ai/leaderboards/models?open_weigh...

This is perfect. Thanks.

Re: Qwen3-4B-Thinking-2507

#50

Earlier quoted context omitted.

China has a business model where you can lose money and it doesn't matter. The state's modus operandi is just fund things until the leader changes his mind about it. This is why the Chinese labs are so open, they don't ever need to make a profit, they just need to make good AI.

Sure. The US isn't morally any better on this though. VC backed companies and the "get users, figure out profitability later" strategy that have made outside countries unable to compete have been enabled by the US government tweaking interest rates and the supply value of the dollar. And the US backs all sorts of unprofitable things through grants, contracts, and bail outs. The US government just hasn't yet found a r…

There isn't really question of morality here, China simply operates different than the US. US companies will need to provide a return to investors and want a secret competitive edge. Chinese companies will need to provide the state with AI, money isn't a concern.

As for why the US gets all the VC love, it's because the US has an extremely friendly business environment. People leave their home countries to start businesses in the US because of it. Europe has done a pretty good job suffocating their tech industry in comparison.

The US government only plays a relatively minor role in this too, because unlike China the US is not a planned economy.

Post reply on HN