Qwen3: Think deeper, act faster
qwenlm.github.io
Qwen3: Think deeper, act faster
1–10 of 412 posts
Re: Qwen3: Think deeper, act faster
#2Not even going into performance, need to test first. But what a stellar release just for attention to all these peripheral details alone. This should be the standard for major release, instead of whatever Meta was doing with Llama 4 (hope Meta can surprise us at LlamaCon tomorrow though).
Re: Qwen3: Think deeper, act faster
#3Re: Qwen3: Think deeper, act faster
#4They have got pretty good documentation too[1]. And Looks like we have day 1 support for all major inference stacks, plus so many size choices. Quants are also up because they have already worked with many community quant makers. Not even going into performance, need to test first. But what a stellar release just for attention to all these peripheral details alone. This should be the standard for major release, inste…
I’m curious, who are the community quant makers?
Re: Qwen3: Think deeper, act faster
#5As this is in trillions, where does this amount of material come from?
Re: Qwen3: Think deeper, act faster
#6Re: Qwen3: Think deeper, act faster
#7They have got pretty good documentation too[1]. And Looks like we have day 1 support for all major inference stacks, plus so many size choices. Quants are also up because they have already worked with many community quant makers. Not even going into performance, need to test first. But what a stellar release just for attention to all these peripheral details alone. This should be the standard for major release, inste…
Re: Qwen3: Think deeper, act faster
#8> The pre-training process consists of three stages. In the first stage (S1), the model was pretrained on over 30 trillion tokens with a context length of 4K tokens. This stage provided the model with basic language skills and general knowledge. As this is in trillions, where does this amount of material come from?
wonder at what price
Re: Qwen3: Think deeper, act faster
#9Probably one of the best parts of this is MCP support baked in. Open source models have generally struggled with being agentic, and it looks like Qwen might break this pattern. The Aider bench score is also pretty good, although not nearly as good as Gemini 2.5 Pro.
I like gemini 2.5 pro a lot bc its fast af but it struggles some times when context is half used to effectively use tools and make edits and breaks a lot of shit (on cursor)
Re: Qwen3: Think deeper, act faster
#10They have got pretty good documentation too[1]. And Looks like we have day 1 support for all major inference stacks, plus so many size choices. Quants are also up because they have already worked with many community quant makers. Not even going into performance, need to test first. But what a stellar release just for attention to all these peripheral details alone. This should be the standard for major release, inste…
Well, the link to huggingface is broken at the moment.