Hy3
hy.tencent.com
Hy3
1–10 of 125 posts
Re: Hy3
#2Re: Hy3
#3Not really at gpt 5.5 tier though, and probably below glm 5.2...
But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model.
Edited: gpt-5.4-mini not the base gpt-5.4
Re: Hy3
#4Re: Hy3
#5Re: Hy3
#6This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
Re: Hy3
#7Re: Hy3
#8This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
hardly, its still quite big unless by "local" you mean people that spend many thousands on rigs :)
Re: Hy3
#9DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Re: Hy3
#10This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.
I looked into DeepSeek's architecture a little bit and the main focus was how can we save as much money as possible. They did a lot of cost cutting with the attention mechanisms. This allowed them to offer an insanely cheap price even on massive contexts, but seems to have come at the cost of performance?
At least, that's my guess, when I see smaller models costing more and outperforming, I think, "they must have denser attention?"