Live data from Hacker News

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

bloomberg.com

91–100 of 155 posts

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#92

I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had t…

> One of them was running the same bash command about a thousand times.

An amusing thought of returning to your workstation to find it as an obsidian block after it gets stuck executing "dd" thousand times.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#93
Ox Alpha has been running on auto-pilot for the past 5 days on various experiments.

Very impressive model.

Here are some examples, open-source documented and the data available in HF datasets:

https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone

https://openzot.github.io/arcade/ - https://github.com/openzot/arcade

https://openzot.github.io/machinery/ - https://github.com/openzot/machinery

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#94

Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?

omp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#95
post #37

There's a lot of brand confusion among the Chinese models right now. Kimi, Qwen, GLM, Z.ai, Ox. We might know the difference (or I should say, someone does because I'm losing track already) but these models have no chance at end user penetration and loyalty until there's a single focused survivor. It took me a year talking about it until my wife knew that ChatGPT and Gemini are two different things. PS: some replies,…

The bubbling froth at the open edge is getting user adopted at a crazy pace, by the early adopter persona trying them all within hours to days. This persona loves taking apart and putting together novel things, and telling others. Fast follower persona clusters around emerging zeitgeist across the tellings. At the moment, arguably that's mostly Qwen for everyday hobbyists, and GLM for those that can run 512GB to 1.5T…

In raw numbers of humans... "the early majority" surely would be those that use ChatGPT or Gemini (aka Google) and pay between $0 and $20 a month?

I would be surprised if the specialist that knows that various Chinese models exist and/or that a user might choose a harness and model separately are a "majority" even of the early variety... in terms of revenue, humans, tokens, or any metric.

(Happy to be proven wrong)

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#96

I had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had t…

Were you using the full model or a quantized version, and what harness/configuration were you using?

It sounds like you were using a quant model.

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#97

Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?

omp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics.

to be fair, OMP/Pi is also just a better harness. e.g. https://www.databricks.com/blog/benchmarking-coding-agents-d...

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#98
post #50

Earlier quoted context omitted.

What would constitute evidence in your opinion?

How about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting?

"Black-Box On-Policy Distillation of Large Language Models", Microsoft Research, https://aka.ms/GAD-project

> 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust solution for black-box LLM distillation.'

No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5?

Re: Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

#99

I’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly differentiating.

Definitely agree. If it is small (eg. Qwen 3.8 28b or gpt-oss-120) then this might be amazing. If it is anywhere near Kimi K3 it would need to have some other differentiating factor than intelligence.
Post reply on HN