Live data from Hacker News

Hy3

hy.tencent.com

21–30 of 125 posts

Re: Hy3

#21

Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.

That's a 2-bit quant of DS4 flash. You're probably better off running Qwen3.6-27B at Q8.

I think its good advice to test both on your own evals for sure, but the MoE parameters are already natively FP4 in ds4. Dropping to 2bpw isn't as big of a loss as it seems (and as corroborated by antirez's work).

Its also only 13B active, so your decode speed would be nearly 2x that of Qwen3.6-27B. So there are other latent benefits as well.

Re: Hy3

#22
post #15

Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization. DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.

Hy3 lacks the DSv4 architecture's KV Cache efficiency. Whereas I can run DSv4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 (quantized to FP4), there is only room for ~130K tokens of KV cache.

Lower context window notwithstanding, Hy3's coding benchmarks hold their own against DeepSeek v4 Pro & MiMo v2.5 Pro. That's quite something for a model priced like DeepSeek v4 Flash & MiMo v2.5 (for non-cached tokens), which are 3x cheaper than their respective Pro variants.

Re: Hy3

#23
post #3

I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap. Not really at gpt 5.5 tier though, and probably below glm 5.2... But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model. Edited: gpt-5.4-mini not the base…

I think you’ve got the models wrong…gpt-5.4? I doubt there is any open source mode matching it. Maybe in a year

Yeah I meant gpt-5.4-mini, but GLM 5.2 is pretty close to gpt-5.4 base, and much better than it when it comes to design stuff.

Re: Hy3

#24
A month ago I wrote a blog post about how Hy3 was topping the OpenRouter rankings despite no one talking about it: https://news.ycombinator.com/item?id=48317294

As of today, it has fallen to 8/9th on the rankings. I don't see a reason where you would use this model over competitors. However, price economics are bit confusing, as currently the effective input price of Hy3 via OpenRouter is now the same as DeepSeek-hosted DeepSeek Flash V4.

https://openrouter.ai/tencent/hy3-preview

https://openrouter.ai/deepseek/deepseek-v4-flash

Re: Hy3

#25
post #3

I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap. Not really at gpt 5.5 tier though, and probably below glm 5.2... But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model. Edited: gpt-5.4-mini not the base…

I think you’ve got the models wrong…gpt-5.4? I doubt there is any open source mode matching it. Maybe in a year

GLM 5.2 already matches GPT-5.4 easily.

Re: Hy3

#26

Earlier quoted context omitted.

DS4-Flash is not only "significantly" smaller, it will also benefit from a lot more speed thanks to DSpark

299B for Hy3 vs 284B* for Flash Edit: fixed, got bad info

flash is 284b isnt it? https://artificialanalysis.ai/models/deepseek-v4-flash

Re: Hy3

#27
post #15

Earlier quoted context omitted.

Hy3 lacks the DSv4 architecture's KV Cache efficiency. Whereas I can run DSv4 Flash on a pair of DGX Sparks and have enough memory left over for 3M tokens of KV cache, with Hy3 (quantized to FP4), there is only room for ~130K tokens of KV cache.

Lower context window notwithstanding, Hy3's coding benchmarks hold their own against DeepSeek v4 Pro & MiMo v2.5 Pro . That's quite something for a model priced like DeepSeek v4 Flash & MiMo v2.5 (for non-cached tokens), which are 3x cheaper than their respective Pro variants.

It's impressive indeed. I would also expect the next checkpoint of DSv4 Flash to come in somewhere at this level (DeepSeek has had over 2 months to continue training since it released).

It's exciting that the open models continue to get better and more efficient across the board!

Re: Hy3

#29

Earlier quoted context omitted.

299B for Hy3 vs 284B* for Flash Edit: fixed, got bad info

flash is 284b isnt it? https://artificialanalysis.ai/models/deepseek-v4-flash

Oh, it is. I was looking at the Huggingface repo which listed the lower number at the top of the page, looks like that's wrong.

Re: Hy3

#30

Quite interesting to see them and Meta and others release before OpenAI supposedly is to release GPT 5.6 today, would it be better to release it before or after? Calm before the storm type of thing?

[deleted]
Post reply on HN