Live data from Hacker News

Qwen3: Think deeper, act faster

qwenlm.github.io

111–120 of 412 posts

Re: Qwen3: Think deeper, act faster

#111
post #42

Earlier quoted context omitted.

The avoiding talking part is more on the Frontend level censorship I think. It doesn't censor on API

This is NOT true. At least on the 1.5B version model on my local machine. It blocks answers when using offline mode. Perplexity has an uncensored a version, but don't thing it is open on how they did it.

Here's a blog post on Perplexity's R1 1776, which they post-trained

https://www.perplexity.ai/hub/blog/open-sourcing-r1-1776

Re: Qwen3: Think deeper, act faster

#112

Earlier quoted context omitted.

4 bit is absolutely fine . I know this is crazy to here because the big iron folks still debate 16 vs 32 and 8 vs 16 is near verboten in public conversation. I contribute to llama.cpp and have seen many many efforts to measure evaluation perf of various quants, and no matter which way it was sliced (ranging from subjective volunteers doing A/B voting on responses over months, to objective object perplexity loss) Q4 i…

It's incredibly niche, but Gemma 3 27b can recognize a number of popular video game characters even in novel fanart (I was a little surprised at that when messing around with its vision). But the Q4 quants, even with QAT, are very likely to name a random wrong character from within the same franchise, even when Q8 quants name the correct character. Niche of a niche, but just kind of interesting how the quantization j…

Vision models do degrade more with quantization. https://unsloth.ai/blog/dynamic-4bit

Re: Qwen3: Think deeper, act faster

#114
post #38

Something that interests me about the Qwen and DeepSeek models is that they have presumably been trained to fit the worldview enforced by the CCP, for things like avoiding talking about Tiananmen Square - but we've had access to a range of Qwen/DeepSeek models for well over a year at this point and to my knowledge this assumed bias hasn't actually resulted in any documented problems from people using the models. Asid…

In my limited experience, models like Llama and Gemma are far more censored than Qwen and Deepseek.

Try to ask any model about Israel and Hamas

Re: Qwen3: Think deeper, act faster

#117
post #98

I have a small physics-based problem I pose to LLMs. It's tricky for humans as well, and all LLMs I've tried (GPT o3, Claude 3.7, Gemini 2.5 Pro) fail to answer correctly. If I ask them to explain their answer, they do get it eventually, but none get it right the first time. Qwen3 with max thinking got it even more wrong than the rest, for what it's worth.

Hi, I'm starting an evals company, would love to have you as an advisor!

Not OP, but what exactly do I need to do.

I'll do it for cheap if you'll let me work remote from outside the states.

Re: Qwen3: Think deeper, act faster

#118
The benchmark results are so incredibly good they are hard to believe. A 30B model that's competitive with Gemini 2.5 Pro and way better than Gemma 27B?

Update: I tested "ollama run qwen3:30b" (the MoE) locally and while it thought much it wasn't that smart. After 3 follow up questions it ended up in an infinite loop.

I just tried again, and it ended up in an infinite loop immediately, just a single prompt, no follow-up: "Write a Python script to build a Fitch parsimony tree by stepwise addition. Take a Fasta alignment as input and produce a nwk string as outpput."

Update 2: The dense one "ollama run qwen3:32b" is much better (albeit slower of course). It still keeps on thinking for what feels like forever until it misremembers the initial prompt.

Re: Qwen3: Think deeper, act faster

#119

Earlier quoted context omitted.

Only in the same way that the plural of 'opinion' is 'fact' ;)

Except, very literally, data is a collection of single points (ie what we call "anecdotes").

Except that the plural of anecdotes is definitely not data, because without controlling for confounding variables and sampling biases, you will get garbage.

Re: Qwen3: Think deeper, act faster

#120

Earlier quoted context omitted.

Hi, I'm starting an evals company, would love to have you as an advisor!

Not OP, but what exactly do I need to do. I'll do it for cheap if you'll let me work remote from outside the states.

I believe they're kidding, playing on "my singular question isn't answered correctly"
Post reply on HN