Live data from Hacker News

Hy3

hy.tencent.com

1–10 of 125 posts

Re: Hy3

#3
I tried out the model it's pretty great, better than ~~gpt5.4~~ gpt-5.4-mini perhaps, atleast close enough to sonnet 5 in performance that I didn't notice much of a gap.

Not really at gpt 5.5 tier though, and probably below glm 5.2...

But most of all it just works for me for most things I tried and it's exceedingly cheap so there is no reason not to use it, if you need a foss model.

Edited: gpt-5.4-mini not the base gpt-5.4

Re: Hy3

#4
This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.

Re: Hy3

#6
post #4

This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.

hardly, its still quite big unless by "local" you mean people that spend many thousands on rigs :)

Re: Hy3

#7
It's a very good model for this size and price. I tried it with a couple of small tasks - just an year ago this would be the level of the leading models.

Re: Hy3

#8
post #4

This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.

hardly, its still quite big unless by "local" you mean people that spend many thousands on rigs :)

Yeah i shouldve been more clear, a model of this size could run on 2 dgx sparks so out of the range of a lot of the typical consumer sure, but I think there is definitely a market for that size

Re: Hy3

#9
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization.

DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.

Re: Hy3

#10
post #4

This model is shockingly small for how capable it is. its a little bit bigger than deepseekV4 flash but around as capable if not more on some benchmarks than V4 pro, i wouldnt be surprised if this becomes a popular local model.

I've been wondering about that. GLM-5.2 is also half the size of DeepSeek V4 Pro. (But costs roughly twice as much.)

I looked into DeepSeek's architecture a little bit and the main focus was how can we save as much money as possible. They did a lot of cost cutting with the attention mechanisms. This allowed them to offer an insanely cheap price even on massive contexts, but seems to have come at the cost of performance?

At least, that's my guess, when I see smaller models costing more and outperforming, I think, "they must have denser attention?"

Post reply on HN