Live data from Hacker News

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

bloomberg.com

21–30 of 135 posts

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#21
post #13

Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?

the 2nd website is not official, just something someone slopped together for some reason.

I have plenty of these websites, I can’t understand why someone is doing this.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#22

Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?

the outperform Fable was a mid (not completed) benchmark run. Real results were lower.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#23

Earlier quoted context omitted.

GLM 5.3 was a great model, so this would be strange to release a regressed model

It's probably GLM 5.3 flash, so weaker but cheaper.

With vision on top

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#24
post #6

I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.

Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.

2. There was a new checkpoint. Official.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#25
post #6

I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.

Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.

3. Deployment problems unrelated to the weights causing degraded performance

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#28
post #6

I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.

It's logical to serve the best version (quant) of the model at the beginning so that users keep testing it. It is also reasonable to think that the developer of the model tried to test various quant levels by gradually degrading the model's capabilities.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#30
post #13

Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?

the 2nd website is not official, just something someone slopped together for some reason.

Even if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities.

The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)

Post reply on HN