This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
I would.
11–20 of 125 posts
This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
I would.
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Whereas I can run DSv4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 (quantized to FP4), there is only room for ~130K tokens of KV cache.
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
DS4-Flash is not only "significantly" smaller, it will also benefit from a lot more speed thanks to DSpark
Edit: fixed, got bad info
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
That's a 2-bit quant of DS4 flash. You're probably better off running Qwen3.6-27B at Q8.
I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap. Not really at gpt 5.5 tier though, and probably below glm 5.2... But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model. Edited: gpt-5.4-mini not the base…
I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap. Not really at gpt 5.5 tier though, and probably below glm 5.2... But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model. Edited: gpt-5.4-mini not the base…
GPT5.4 xhigh DeepSWE - 52%
A lot of contaminated benchmarks in the blog post about Hy3, needs real testing though I have a distinct feeling it's benchmaxxed like a lot of Chinese models.