> The pre-training process consists of three stages. In the first stage (S1), the model was pretrained on over 30 trillion tokens with a context length of 4K tokens. This stage provided the model with basic language skills and general knowledge. As this is in trillions, where does this amount of material come from?
Qwen3: Think deeper, act faster
141–150 of 412 posts
Re: Qwen3: Think deeper, act faster
#142Something that interests me about the Qwen and DeepSeek models is that they have presumably been trained to fit the worldview enforced by the CCP, for things like avoiding talking about Tiananmen Square - but we've had access to a range of Qwen/DeepSeek models for well over a year at this point and to my knowledge this assumed bias hasn't actually resulted in any documented problems from people using the models. Asid…
The avoiding talking part is more on the Frontend level censorship I think. It doesn't censor on API
Deepseek v3: Taiwan is not a country; it is an inalienable part of China's territory. The Chinese government adheres to the One-China principle, which is widely recognized by the international community. (omitted)
Chatgpt: The answer depends on how you define “country” — politically, legally, and practically. In practice: Taiwan functions like a country. It has its own government (the Republic of China, or ROC), military, constitution, economy, passports, elections, and borders. (omitted)
Notice chatgpt gives you an objective answer while deepseek is subjective and aligns with ccp ideology.
Re: Qwen3: Think deeper, act faster
#143Re: Qwen3: Think deeper, act faster
#144Earlier quoted context omitted.
I don't really want it added to the training set, but eh. Here you go: > Assume I have a 3D printer that's currently printing, and I pause the print. What expends more energy, keeping the hotend at some temperature above room temperature and heating it up the rest of the way when I want to use it, or turning it completely off and then heat it all the way when I need it? Is there an amount of time beyond which the ans…
Some calculation around heat loss and required heat expenditure to reheat per material or something?
Re: Qwen3: Think deeper, act faster
#145I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.
> I put a coin in a cup and slam it upside-down on a glass table. I can't see the coin because the cup is over it. I slide a mirror under the table and see heads. What will I see if I take the cup (and the mirror) away?
Re: Qwen3: Think deeper, act faster
#146Something that interests me about the Qwen and DeepSeek models is that they have presumably been trained to fit the worldview enforced by the CCP, for things like avoiding talking about Tiananmen Square - but we've had access to a range of Qwen/DeepSeek models for well over a year at this point and to my knowledge this assumed bias hasn't actually resulted in any documented problems from people using the models. Asid…
Re: Qwen3: Think deeper, act faster
#147China is doing a great job raising doubt about any lead the major US labs may still have. This is solid progress across the board. The new battlefront may be to take reasoning to the level of abstraction and creativity to handle math problems without a numerical answer (for ex: https://arxiv.org/pdf/2503.21934 ). I suspect that kind of ability will generalize well to other areas and be a significant step toward human…
Re: Qwen3: Think deeper, act faster
#148Earlier quoted context omitted.
Can you please share the problem?
I don't really want it added to the training set, but eh. Here you go: > Assume I have a 3D printer that's currently printing, and I pause the print. What expends more energy, keeping the hotend at some temperature above room temperature and heating it up the rest of the way when I want to use it, or turning it completely off and then heat it all the way when I need it? Is there an amount of time beyond which the ans…
Re: Qwen3: Think deeper, act faster
#149Earlier quoted context omitted.
There are a lot of variables here such as your hardware's memory bandwidth, speed at which at processes tensors etc. A basic thing to remember: Any given dense model would require X GB of memory at 8-bit quantization, where X is the number of params (of course I am simplifying a little by not counting context size). Quantization is just 'precision' of the model, 8-bit generally works really well. Generally speaking,…
4 bit is absolutely fine . I know this is crazy to here because the big iron folks still debate 16 vs 32 and 8 vs 16 is near verboten in public conversation. I contribute to llama.cpp and have seen many many efforts to measure evaluation perf of various quants, and no matter which way it was sliced (ranging from subjective volunteers doing A/B voting on responses over months, to objective object perplexity loss) Q4 i…
For larger models.
For smaller models, about 12B and below, there is a very noticeable degradation.
At least that's my experience generating answers to the same questions across several local models like Llama 3.2, Granite 3.1, Gemma2 etc and comparing Q4 against Q8 for each.
The smaller Q4 variants can be quite useful, but they consistently struggle more with prompt adherence and recollection especially.
Like if you tell it to generate some code without explaining the generated code, a smaller Q4 is significantly more likely to explain the code regardless, compared to Q8 or better.
Re: Qwen3: Think deeper, act faster
#150I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.
You really had me until the last half of the last sentence.