Earlier quoted context omitted.
Qwen 35B isn't even remotely close to the big models. It's just people over hyping small models. Ignore the benchmarks they are almost meaningless. If you want something comparable you need the trillion parameter open models like deepseek.
Number of parameters doesn't make the model smarter, it just makes it know more stuff out of the box. At some point there's diminishing returns and your coding LLM performs worse because you encoded useless stuff like Pokemon combinations or languages you don't speak into its parameter space. The "smartness" of the model comes from RLHF post-training, which is orthogonal to model size. Also, if you're using an agenti…
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
101–110 of 116 posts
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#102It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.
I'd love to have multiple large VRAM GPUs but I can't justify the costs when I have plenty of other more important things to spend that money on.
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#103It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.
2) I think it’s important to recognize that one of the things models are good for is astroturfing, and any given conversation you see may be direct or secondary effects of that (among other marketing).
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#104That's almost exactly my setup and I'm very happy with its performance. I noticed recently that I started to prefer my local Qwen3.6 35B A3B and pi agent over Claude Code. Both fail at different tasks, and Qwen more so than Claude. But the way Qwen fails is much more straightforward. In writing tasks Qwens hallucinations and bullshitting are much easier to spot because it doesn't have the sleek vocabulary and wordsmi…
>In writing tasks Qwens hallucinations and bullshitting are much easier to spot because it doesn't have the sleek vocabulary and wordsmithing skills to disguise its ignorance. Can't wait until we just remove the language from the LLMs for accuracy and efficiency
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#105Earlier quoted context omitted.
I'm running Qwen3.6-35B-A3B on a very ordinary desktop PC (32GB DDR5, 8GB Radeon 6600XT) and getting a useful 15-20 tok/sec out of it. The MoE architecture and auto offloading from system to VRAM is just fantastic. Unsloth Q4_K_XL. The Qwen3.6-27B is unbearably slow as it doesn't fit in VRAM, though, i think the MoE is very easy to run. It is also extremely nice that you can just `apt install llama.cpp libggml0-backe…
I wonder what parent poster means with „useful” and what he actually tried? Feels like he was just comparing some benchmarks. Yesterday I downloaded Gemma4-26B with Ollama on quite rusty desktop with 1070 8gb and 32gb of ram and Core i5-9400. I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. Earlier li…
"Useful" as in "has a use that isn't just for show". It takes me two seconds to read a photo of a water meter. Having an LLM read it for me in 3 minutes isn't useful. Similarly small models are capable of tool use (e.g. web searches) but their synthesis leaves much to be desired. As an example I'd ask some small models to find examples of products with specific characteristics and they'd come back with only one or two because they discounted other possibilities incorrectly by reasoning themselves out of it.
> Feels like he was just comparing some benchmarks.
On what do you base this assertion?
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#106It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.
14 year old me is mortified at this community. Same here. There has to be someplace like this that's managed to cultivate a better crowd, but I'll be darned if I can find it.
Hn really is a bit of an echo chamber, not that diverse opinions aren’t expressed, but the folks that voice opinions that don’t align with a specific set of values aren’t very well received here.
I’ll also say that this place has shaped my values, including making me change my opinions on things I severely disagreed with at the time. I’ve also said a lot of shit on here I wish I could wipe out.
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#107Earlier quoted context omitted.
14 year old me is mortified at this community. Same here. There has to be someplace like this that's managed to cultivate a better crowd, but I'll be darned if I can find it.
This place is probably the best you’ll find. I actually found hn from a site called meta-somethingOrOther, a long time ago, and that site is probably the closest I can think of. Hn really is a bit of an echo chamber, not that diverse opinions aren’t expressed, but the folks that voice opinions that don’t align with a specific set of values aren’t very well received here. I’ll also say that this place has shaped my va…
I feel the same here when "hackers" defend copyright maximalism, try to rehabilitate the Luddites, and argue that the Federal government should aggressively regulate AI models. Basically exhibiting both proud ignorance of history and reckless disregard for the future, all in one breath.
There are so many other places for that. So very, very many. Why do they come here? I spend time in those places as well, but I generally STFU when I have nothing to contribute, or when my core values conflict with their community charter.
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#108Earlier quoted context omitted.
I wonder what parent poster means with „useful” and what he actually tried? Feels like he was just comparing some benchmarks. Yesterday I downloaded Gemma4-26B with Ollama on quite rusty desktop with 1070 8gb and 32gb of ram and Core i5-9400. I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. Earlier li…
> I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. "Useful" as in "has a use that isn't just for show". It takes me two seconds to read a photo of a water meter. Having an LLM read it for me in 3 minutes isn't useful. Similarly small models are capable of tool use (e.g. web searches) but their synthes…
Mostly use of this expression.
I don’t get agent to read the meter for me - I can do that when I take the photo.
I send the photo to a bot that ingests photos from me and stores readings for me with date and time so later I can ask „what was last reading” or what was the usage between x and y dates”, without me having to make a perfect photo, without me having to dabble with OpenCV.
Even if it takes 30mins it is still useful for me.
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#109The recommended values for Qwen 3.6 in thinking mode is `--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00`, and `--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00` for coding/tool calling tasks, and for non-thinking, `--temp 0.7 -top-p 0.8 --top-k 20 --presence-penalty 1.5 --min-p 0.00`. The options listed are none of these. Also, the recommended Qwen MTP settings are `--spec-type draft-mtp --spec-draft-n-max 2`. 3 is n…
> You can also add `ngram-mod`, but after `draft-mtp` It looks like there's a hardcoded preference, CLI order is not important. (speculative.cpp:1322-1381): common_get_enabled_speculative_configs converts the types vector to a bitmask (order-independent). Then configs are added in a hardcoded priority order: ngram-simple ngram-map-k ngram-map-k4v ngram-mod ngram-cache draft-simple draft-eagle3 draft-mtp (speculative.…
Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
#110Earlier quoted context omitted.
Yeah that’s definitely the smarter buy if you want to just have models running quickly. But the cost of 2 p150 and a 4090 was The main issue is the immature software, and somewhat baroque way of writing kernels. Please, buy one and join us.
Were you able to connect the two P150 using the qsfp-dd cable? They only sell 4x and 8x topologies so I’m curious if that worked for you. Are you able to run them tensor parallel?
I’m not even using a 800G cable since they are expensive and I don’t think I need the bandwidth, opting for 400G instead. This just needs a config change for the number of Ethernet links it uses internally. (Apparently these cables are just many 200G links put together.)