Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
the 2nd website is not official, just something someone slopped together for some reason.
Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
21–30 of 142 posts
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#22Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#23Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#24I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#25I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
Two potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#26Unfortunately I can't find sources other than this for now but this seems to be legit.
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#27Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#28I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#29Funny how all china companies are expected to release weights by default
Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
#30Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
the 2nd website is not official, just something someone slopped together for some reason.
The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)