Live data from Hacker News

RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

imil.net

101–110 of 116 posts

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#101
post #88

Earlier quoted context omitted.

Qwen 35B isn't even remotely close to the big models. It's just people over hyping small models. Ignore the benchmarks they are almost meaningless. If you want something comparable you need the trillion parameter open models like deepseek.

Number of parameters doesn't make the model smarter, it just makes it know more stuff out of the box. At some point there's diminishing returns and your coding LLM performs worse because you encoded useless stuff like Pokemon combinations or languages you don't speak into its parameter space. The "smartness" of the model comes from RLHF post-training, which is orthogonal to model size. Also, if you're using an agenti…

[deleted]

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#102

It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.

The problem is accessibility. GPUs have gotten expensive and in some cases hard to get. You also need the supporting hardware to use them.

I'd love to have multiple large VRAM GPUs but I can't justify the costs when I have plenty of other more important things to spend that money on.

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#103

It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.

1) Different people might optimize for different things. There are people calculating that expensive hardware plus cheap rentals means owning isn’t optimizing, but there are people making choices that fit your preferences too.

2) I think it’s important to recognize that one of the things models are good for is astroturfing, and any given conversation you see may be direct or secondary effects of that (among other marketing).

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#104
post #11

That's almost exactly my setup and I'm very happy with its performance. I noticed recently that I started to prefer my local Qwen3.6 35B A3B and pi agent over Claude Code. Both fail at different tasks, and Qwen more so than Claude. But the way Qwen fails is much more straightforward. In writing tasks Qwens hallucinations and bullshitting are much easier to spot because it doesn't have the sleek vocabulary and wordsmi…

>In writing tasks Qwens hallucinations and bullshitting are much easier to spot because it doesn't have the sleek vocabulary and wordsmithing skills to disguise its ignorance. Can't wait until we just remove the language from the LLMs for accuracy and efficiency

Imagine how accurate you could be if you could circumvent the coding agent and just type the source code directly into an editor all by yourself o_O Like a write_file skill but for humans!!!

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#105
post #79
post #62

Earlier quoted context omitted.

I'm running Qwen3.6-35B-A3B on a very ordinary desktop PC (32GB DDR5, 8GB Radeon 6600XT) and getting a useful 15-20 tok/sec out of it. The MoE architecture and auto offloading from system to VRAM is just fantastic. Unsloth Q4_K_XL. The Qwen3.6-27B is unbearably slow as it doesn't fit in VRAM, though, i think the MoE is very easy to run. It is also extremely nice that you can just `apt install llama.cpp libggml0-backe…

I wonder what parent poster means with „useful” and what he actually tried? Feels like he was just comparing some benchmarks. Yesterday I downloaded Gemma4-26B with Ollama on quite rusty desktop with 1070 8gb and 32gb of ram and Core i5-9400. I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. Earlier li…

> I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct.

"Useful" as in "has a use that isn't just for show". It takes me two seconds to read a photo of a water meter. Having an LLM read it for me in 3 minutes isn't useful. Similarly small models are capable of tool use (e.g. web searches) but their synthesis leaves much to be desired. As an example I'd ask some small models to find examples of products with specific characteristics and they'd come back with only one or two because they discounted other possibilities incorrectly by reasoning themselves out of it.

> Feels like he was just comparing some benchmarks.

On what do you base this assertion?

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#106

It is absolutely mind blowing to see some of the responses here. Open source, run-your-own, pay for nothing, we’re-all-nerds-that-buy-the-hardware-anyways ethos seems basically dead. I guess I’m getting old. I own two 16gb cards and I use them for models, for gpu-pasthru for gaming, 3d model rendering, etc. 14 year old me is mortified at this community.

14 year old me is mortified at this community. Same here. There has to be someplace like this that's managed to cultivate a better crowd, but I'll be darned if I can find it.

This place is probably the best you’ll find. I actually found hn from a site called meta-somethingOrOther, a long time ago, and that site is probably the closest I can think of.

Hn really is a bit of an echo chamber, not that diverse opinions aren’t expressed, but the folks that voice opinions that don’t align with a specific set of values aren’t very well received here.

I’ll also say that this place has shaped my values, including making me change my opinions on things I severely disagreed with at the time. I’ve also said a lot of shit on here I wish I could wipe out.

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#107

Earlier quoted context omitted.

14 year old me is mortified at this community. Same here. There has to be someplace like this that's managed to cultivate a better crowd, but I'll be darned if I can find it.

This place is probably the best you’ll find. I actually found hn from a site called meta-somethingOrOther, a long time ago, and that site is probably the closest I can think of. Hn really is a bit of an echo chamber, not that diverse opinions aren’t expressed, but the folks that voice opinions that don’t align with a specific set of values aren’t very well received here. I’ll also say that this place has shaped my va…

I'm fine with diverse opinions, as long as they're not too diverse... and yes, there is such a thing as "too diverse." If I were to barge into an Amish town meeting and harangue them about how they should be using Qwen 3.6 27B Q8 to plan their crop rotation schedule, I would soon find myself heading out of town facing south on a northbound mule. And that's OK.

I feel the same here when "hackers" defend copyright maximalism, try to rehabilitate the Luddites, and argue that the Federal government should aggressively regulate AI models. Basically exhibiting both proud ignorance of history and reckless disregard for the future, all in one breath.

There are so many other places for that. So very, very many. Why do they come here? I spend time in those places as well, but I generally STFU when I have nothing to contribute, or when my core values conflict with their community charter.

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#108
post #79

Earlier quoted context omitted.

I wonder what parent poster means with „useful” and what he actually tried? Feels like he was just comparing some benchmarks. Yesterday I downloaded Gemma4-26B with Ollama on quite rusty desktop with 1070 8gb and 32gb of ram and Core i5-9400. I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. Earlier li…

> I drop photo of my water meter and tell it to read the value and serial number. It was far from instant but it was also easily under 3 minutes and result was correct. "Useful" as in "has a use that isn't just for show". It takes me two seconds to read a photo of a water meter. Having an LLM read it for me in 3 minutes isn't useful. Similarly small models are capable of tool use (e.g. web searches) but their synthes…

trying to scheme a way

Mostly use of this expression.

I don’t get agent to read the meter for me - I can do that when I take the photo.

I send the photo to a bot that ingests photos from me and stores readings for me with date and time so later I can ask „what was last reading” or what was the usage between x and y dates”, without me having to make a perfect photo, without me having to dabble with OpenCV.

Even if it takes 30mins it is still useful for me.

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#109
post #93

The recommended values for Qwen 3.6 in thinking mode is `--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00`, and `--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00` for coding/tool calling tasks, and for non-thinking, `--temp 0.7 -top-p 0.8 --top-k 20 --presence-penalty 1.5 --min-p 0.00`. The options listed are none of these. Also, the recommended Qwen MTP settings are `--spec-type draft-mtp --spec-draft-n-max 2`. 3 is n…

> You can also add `ngram-mod`, but after `draft-mtp` It looks like there's a hardcoded preference, CLI order is not important. (speculative.cpp:1322-1381): common_get_enabled_speculative_configs converts the types vector to a bitmask (order-independent). Then configs are added in a hardcoded priority order: ngram-simple ngram-map-k ngram-map-k4v ngram-mod ngram-cache draft-simple draft-eagle3 draft-mtp (speculative.…

Interesting.

Re: RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

#110
post #100
post #81

Earlier quoted context omitted.

Yeah that’s definitely the smarter buy if you want to just have models running quickly. But the cost of 2 p150 and a 4090 was The main issue is the immature software, and somewhat baroque way of writing kernels. Please, buy one and join us.

Were you able to connect the two P150 using the qsfp-dd cable? They only sell 4x and 8x topologies so I’m curious if that worked for you. Are you able to run them tensor parallel?

Yeah, I’m doing TP with two cards. The topology is configured based on yaml files, and if you are not using a predefined config you can just create a new config with your topology.

I’m not even using a 800G cable since they are expensive and I don’t think I need the bandwidth, opting for 400G instead. This just needs a config change for the number of Ethernet links it uses internally. (Apparently these cables are just many 200G links put together.)

Post reply on HN