Live data from Hacker News

Unsloth Dynamic 3.0 GGUFs

unsloth.ai

41–50 of 125 posts

Re: Unsloth Dynamic 3.0 GGUFs

#41
post #30
post #12

Cool. Now run TerminalHard and compare to unquantized 27B. KLD of 1%, or similar error metric that multiplies, on 10000 tokens would give accumulated error of 2,000,000%

I don't think you can extrapolate that measurement across multiple sequential draws like that. We presumably are comparing against a single trajectory rather than a tree of trajectories. So once we make the wrong choice and step off of the blessed path, we have no way to assign a ranking to the next token; it's error is undefined. I've seen LLMs self correct in chains of thought ("because of foo and bar, I need to...…

actually, bar is true

but wait, the models constantly go back and forth on these things in their thinking traces, so it is unclear which self correcting is actually correct

Re: Unsloth Dynamic 3.0 GGUFs

#42
post #25
post #12

Cool. Now run TerminalHard and compare to unquantized 27B. KLD of 1%, or similar error metric that multiplies, on 10000 tokens would give accumulated error of 2,000,000%

That's not how that works. Selecting a different token is not inherently erroneous. A correct solution can still be found despite divergence.

KLD isn't how that works either. The truth is in the middle and they aren't showing it.

Re: Unsloth Dynamic 3.0 GGUFs

#43
post #2

Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? impro…

Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed.

As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that.

But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL

Re: Unsloth Dynamic 3.0 GGUFs

#44

Are there benchmarks for the various Qwen3.8-27B quants that actually measure writing code, maybe even with multiple steps? Low KL divergence does not mean much when the model gets stuck in doom loops all the time. I could of course download and test myself, but that would take days with my internet connection.

We made something called Divergence-300 @32 (and later @512) which tests actual inference across 32 tokens on a held out test (Terminal Bench, DeepSWE, Math etc)

We do plan to do larger benchmark suites though!

Re: Unsloth Dynamic 3.0 GGUFs

#45
post #2

Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? impro…

Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed. As someone in the comments said we made a separate Q4_0 MTP if that's helpful so you can use that. But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL

Daniel, question I got the Qwen3.8-27B-UD-Q2_K_XL.gguf from https://huggingface.co/unsloth/Qwen3.8-27B-GGUF?show_file_in... and continue with my testing, but the model quickly felt into a loop of asking the same thing over and over again, I have seen the MOE do that but not the dense ones.

And I had similar experiences when Qwen3.8-27B unsloth images just came out with the full Q8_K_XL, I'm using an AMD setup which has modifications to save to disk the kv, but your (assuming you are part of the unsloth team) for some reason have been giving me similar issues.

I tried https://huggingface.co/mradermacher/Qwen3.8-27B-Uncensored-G... the 8 bit, 6 and 2 bit... the 2 bit almost use the complete KV doing it's thing and didn't loop itself.

It can be something in my setup, there is a very high chance of that, but the previous 3.6 images from qwen, the 27B, the 31A3 and 122 they are all unsloth and did work on my setup without issues...

Again could be my setup... let me know if there is any data I can supply to you to debug if needed.

Re: Unsloth Dynamic 3.0 GGUFs

#46
post #26

Earlier quoted context omitted.

What about do you mean by single threaded? Each token is predicted by using parallel computation on the GPU.

Multiple agents need tokens. Should optimize for that instead of one agent blocking the others.

One agent typically blocks the others on a local device because the GPU is already completely utilized either in terms of memory or compute. You can have true parallelism at home, but you need an absurd amount of resources. It's not a simple threading problem.

Re: Unsloth Dynamic 3.0 GGUFs

#47
post #2

Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? impro…

Q2 quantization is basically giving a capable model a lobotomy. It will not accurately represent how smart or capable something like qwen 3.8 27B in Q8 will be.

Re: Unsloth Dynamic 3.0 GGUFs

#48

"We also made some smaller UD-1bit quants with UD-IQ1_S being 6.2GB (without MTP) which retain around 72% top-1% accuracy yet being 89% smaller" This is crazy! But has anyone tried these lower quants on real projects?

I tried some 1-bit, 2-bit, and bonsai quants against closed eval sets. They were essentially useless for my case. The little errors accumulate and send the whole output off track quickly.

If you had some use case with very small output sequences they could be interesting to try. I think dropping down to a 9B-class model would produce better results for most cases.

Re: Unsloth Dynamic 3.0 GGUFs

#49

Earlier quoted context omitted.

Yes, you can. Ideally though, you want to minimize the number of cards and maximize the amount of memory in each card. Using multiple cards is one of the things that the models and software that Unsloth releases does really well in terms of ease of use and relatively good performance.

How many do you recommend?

I've run 4, 6, and 8. Adding more video cards increases total available memory, so you can load larger models, but there are definitely some drawbacks.

More cards = more communication over PCIe. The prompts don't come in and just get magically split between each card, they move sequentially through them.

Also, with 2x24GB cards you don't really have 48GB of usable memory to load a model, closer to ~42GB + context.

And then there are power concerns, motherboard limitations (PCIe slots and lanes - a lot of motherboards with multiple 16x PCIe slots don't actually have 16x lanes to each of those slots), and more. 8x GPUs are going to easily draw 2000W on their own, if not substantially more. You'll need wiring and a circuit that can support 3000W without a risk of starting a fire in your wall.

For $5k, a single 32GB 5090 might be a better choice for a lot of people versus 4x3090s with 24GB each. It will definitely perform substantially better on smaller 27B models.

For hardware:

A good motherboard with lots of PCIe lanes (7x full 16x PCIe 4.0), DDR4 support, etc:

https://www.asus.com/us/motherboards-components/motherboards...

Add in a 3xxx series Threadripper PRO, 128 or 256GB of DDR4 (going higher becomes really expensive), and a ~1400 watt power supply. You can underpower/undervolt Nvidia cards really easily, and capping them at 250W loses you minimal performance.

Re: Unsloth Dynamic 3.0 GGUFs

#50
It would be nice if unsloth published GGUFs would use a version number or something, because now I have multiple different files on local storage that otherwise have exactly the same name.

"Qwen3.8-27B-UD-Q8_K_XL.gguf" for instance.

The one downloaded at least 4 days ago is a different thing and is NOT the "Dynamic 3.0" GGUF which I am now downloading, which I presume will have a different sha256 checksum?

The unsloth page says dynamic 3.0 is released "today", but I have an older copy of qwen3.8 27B Q8 which I downloaded, if I remember right, at least 4-5 days ago...

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF

Post reply on HN