Earlier quoted context omitted.
Unlike cloud infra in general which offers things like automatic backups, regional redundancy, and effectively unlimited scalability, it seems like the value proposition of cloud LLM gets ever shakier. * Many businesses don't need frontier level intelligence anyway. * It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well? Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries. Plus, you know, the whole economies…
Qwen 3.8 27B is excellent, but it defaults to overthinking things
121–130 of 411 posts
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#122“The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year.” I think that should be the blinking headline - this shows what can be done with consumer hardware.
Tried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#123Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#124Earlier quoted context omitted.
I think not much can run without a dedicated GPU
Till now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed
If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#125But what I did see is that it does overthink a lot.
17GB is Q4 for Qwen3.8. That's quantitized quite a bit.
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#126Earlier quoted context omitted.
I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well? Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries. Plus, you know, the whole economies…
All white collar work will be done by AI soon, there won’t be scaling down, just scaling up. I do food trucks and festivals and I’ve got AI doing so much of my non-meatspace work now that I’m buying $10-$20 a day in tokens. At that rate, hardware starts to look cheap, and I have less of a use case for it than most white collar workers.
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#127Earlier quoted context omitted.
All white collar work will be done by AI soon, there won’t be scaling down, just scaling up. I do food trucks and festivals and I’ve got AI doing so much of my non-meatspace work now that I’m buying $10-$20 a day in tokens. At that rate, hardware starts to look cheap, and I have less of a use case for it than most white collar workers.
Genuine question, what are you spending that on? The $20/month ChatGPT/Codex subscription has largely been enough for me as an IT worker.
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#128Earlier quoted context omitted.
Unlike cloud infra in general which offers things like automatic backups, regional redundancy, and effectively unlimited scalability, it seems like the value proposition of cloud LLM gets ever shakier. * Many businesses don't need frontier level intelligence anyway. * It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
The whole cloud story lies on two aspects: - Hyperscaling “we are going to serve billions of people in our applications”, which is becoming increasing unlikely as regional tech companies become more dominant than than the global one (this one is as much about geopolitics as technology) - Operations is hard, in which case non-frontier models should be increasingly capable. Devops for small-ish deployment is one of the…
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#129Earlier quoted context omitted.
I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well? Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries. Plus, you know, the whole economies…
All white collar work will be done by AI soon, there won’t be scaling down, just scaling up. I do food trucks and festivals and I’ve got AI doing so much of my non-meatspace work now that I’m buying $10-$20 a day in tokens. At that rate, hardware starts to look cheap, and I have less of a use case for it than most white collar workers.
Re: Qwen 3.8 27B is excellent, but it defaults to overthinking things
#130Earlier quoted context omitted.
Unlike cloud infra in general which offers things like automatic backups, regional redundancy, and effectively unlimited scalability, it seems like the value proposition of cloud LLM gets ever shakier. * Many businesses don't need frontier level intelligence anyway. * It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well? Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries. Plus, you know, the whole economies…