Live data from Hacker News

Ox-Alpha Is GLM?

dejan.ai

51–60 of 75 posts

Re: Ox-Alpha Is GLM?

#51
The Ox-Alpha webpage really make it sound like they are trying to hype a model that has nothing particular to show:

"The reasoning model that appeared out of nowhere. Built for code, long-horizon agents, and a million tokens of context. Nobody knows who made it — everyone wants to try it."

Re: Ox-Alpha Is GLM?

#52
post #2

GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so. My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot. MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant…

Z.ai founders were one of the pioneers in MM-LLMs with Cog-VLM years ago, back when LLaVA emerged. I wouldn't be surprised if they added multi-modal capabilities

Re: Ox-Alpha Is GLM?

#53

My strong prediction: this model is around 64B and can run on laptops. Thats the reason behind the hype.

It's fun to imagine that it could be GLM 5.3-Flash. Between GLM 4 and 5, the flagship's total parameters doubled and the active parameters went up 25%. GLM 4.7-Flash was 30B / 3B active. If this model were 60B / 4B active, that sure would hit a sweet, currently empty spot in the lineup of open models.

But I don't know. There are vagueposts on X about this model running on two DGX Sparks. If they did the same scale-up with Air instead of Flash, that would probably be about right.

Re: Ox-Alpha Is GLM?

#54
post #42
post #17

If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?

Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5

    > 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.

Re: Ox-Alpha Is GLM?

#56
post #45

My bet: it's Google running a "new" model based on GLM

Interesting. I like this theory because it explains why some G employees were vague posting about it. But like.. why? Why wouldn’t Google just use Gemma?

> why some G employees were vague posting about it

They were just trolling.

Re: Ox-Alpha Is GLM?

#57

My strong prediction: this model is around 64B and can run on laptops. Thats the reason behind the hype.

> this model is around 64B and can run on laptops.

We're lucky if it'd fit in 1 DGX Spark. Laptops - nah, unless you mean like an M5 Max with 128GB of RAM then maybe.

> Thats the reason behind the hype.

The hype is imagine DeepSeek Flash before the price increase with even better performance. It'd be like unlimited Sonnet.

Re: Ox-Alpha Is GLM?

#58
It is probably from Google and is probably hosted on Vertex AI. Opencode announced that responses from Ox Alpha should be better and soon posted about Vertex eu and us multi region update in their changelog. I could be a Gemini model or a new one based on GLM based on tokenizer. Also the amount of inference it is providing for free is something only google can support with its TPUs. So maybe a GLM based model running on TPUs.

Re: Ox-Alpha Is GLM?

#59
post #58

It is probably from Google and is probably hosted on Vertex AI. Opencode announced that responses from Ox Alpha should be better and soon posted about Vertex eu and us multi region update in their changelog. I could be a Gemini model or a new one based on GLM based on tokenizer. Also the amount of inference it is providing for free is something only google can support with its TPUs. So maybe a GLM based model running…

I've seen a number of people report that it answers near-identically to mainland CN built models on topics related to controversial things the CCP doesn't want to talk about. I'd be extremely surprised if it's a Google model.

Re: Ox-Alpha Is GLM?

#60
post #42

Earlier quoted context omitted.

Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5

> 1 quadrillion tokens per day on Nous portal If you are referring to this number ( https://xcancel.com/NousResearch/status/2090899914700054780 ), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.

I was hitting 429 overloaded regularly with Ox on OpenRouter yesterday, but a lot of that turned out to be problems with my harness. I fixed some bugs, improved the back-off, and I haven't hit a 429 error since (touch wood).

OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day.

https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300

Post reply on HN