These performance numbers look absolutely incredible. The MoE outperforms o1 with 3B active parameters? We're really getting close to the point where local models are good enough to handle practically every task that most people need to get done.
Qwen3: Think deeper, act faster
21–30 of 412 posts
Re: Qwen3: Think deeper, act faster
#22These performance numbers look absolutely incredible. The MoE outperforms o1 with 3B active parameters? We're really getting close to the point where local models are good enough to handle practically every task that most people need to get done.
Re: Qwen3: Think deeper, act faster
#23Re: Qwen3: Think deeper, act faster
#24Until now I found that open weight models were either not as good as their proprietary counterparts or too slow to run locally. This looks like a good balance.
Re: Qwen3: Think deeper, act faster
#25Any news on some viable successor of LLMs that could take us to AGI? As I see they still can't solve some fundamental stuff to make it really work in any scenario (halucinations, reasoning, grounding in reality, updating long-term memory, etc.)
On the other hand, LLM progress feels like bullshit, gaming benchmarks and other problems occured. So either in two years all hail our AGI/AMI (machine intelligence) overlords, or the bubble bursts.
Re: Qwen3: Think deeper, act faster
#26Re: Qwen3: Think deeper, act faster
#27Any news on some viable successor of LLMs that could take us to AGI? As I see they still can't solve some fundamental stuff to make it really work in any scenario (halucinations, reasoning, grounding in reality, updating long-term memory, etc.)
They do improve on literally all of these, at incredible speed and without much sign of slowing down.
Are you asking for a technical innovation that will just get from 0 to perfect AI? That is just not how reality usually works. I don't see why of all things AI should be the exception.
Re: Qwen3: Think deeper, act faster
#28Re: Qwen3: Think deeper, act faster
#29Earlier quoted context omitted.
they have already worked with many community quant makers I’m curious, who are the community quant makers?
I had Unsloth[1] and Bartowski[2] in mind. Both said on Reddit that Qwen had allowed them access to weights before release to ensure smooth sailing. [1] https://huggingface.co/unsloth [2] https://huggingface.co/bartowski
Re: Qwen3: Think deeper, act faster
#30These performance numbers look absolutely incredible. The MoE outperforms o1 with 3B active parameters? We're really getting close to the point where local models are good enough to handle practically every task that most people need to get done.
How do people typically do napkin math to figure out if their machine can “handle” a model?