Qwen 3.8 27B is excellent, but it defaults to overthinking things
101–110 of 411 posts
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#102Complaining about overthinking in xhigh then pointing out output had bugs with thinking turned off seems like it’s missing the obvious compromise?
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#103To me, the amazing thing is that we now have local models that rival the reasoning of high end models from about a year ago. I hope this trend continues.
Unlike cloud infra in general which offers things like automatic backups, regional redundancy, and effectively unlimited scalability, it seems like the value proposition of cloud LLM gets ever shakier. * Many businesses don't need frontier level intelligence anyway. * It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
- Hyperscaling “we are going to serve billions of people in our applications”, which is becoming increasing unlikely as regional tech companies become more dominant than than the global one (this one is as much about geopolitics as technology)
- Operations is hard, in which case non-frontier models should be increasingly capable. Devops for small-ish deployment is one of the few cases where it is hard to clam you need deep expertise and AI can’t do it. Previously, the claim is that you need people specialized in ops, which is expensive. Now…
My prediction is that not just cloud LLM, but cloud business general will have to change. Not yet in the next 5 years, but probably 8-20 years-ish
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#104I do agree that Qwen 3.8 27B is excellent but slow and very token inefficient. My benchmark places it near opus 4.6 and codex 5.3 performance. 3.6 27B couldn't even complete the benchmark. Please see below for details: https://gist.github.com/nharziro/aed0c364ce2f295a493494c6f1b...
Opus 4.6 performance with a local model that can be hosted on consumer hardware is an incredible result!!
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#105Earlier quoted context omitted.
Tried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
I think not much can run without a dedicated GPU
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#106Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#107Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#108Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#109What gets reported is always the benchmark result, but rarely the real-world trade-off made to achieve it. That’s an obvious incentive for the labs, so I think Simon is correctly zeroing in on it. Please continue doing so for models that don’t go too far as much as this release.
Don’t get me wrong, I think it’s amazing what we can get out of smaller models with more reasoning, but we should be super aware how very much not-free it is.
This is a good opportunity to call out models that reason quickly: Meta’s Glimmer seems to be pretty token efficient so far, as do the GPT 5.6s.