"The reasoning model that appeared out of nowhere. Built for code, long-horizon agents, and a million tokens of context. Nobody knows who made it — everyone wants to try it."
Ox-Alpha Is GLM?
51–60 of 75 posts
Re: Ox-Alpha Is GLM?
#52GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so. My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot. MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant…
Re: Ox-Alpha Is GLM?
#53My strong prediction: this model is around 64B and can run on laptops. Thats the reason behind the hype.
But I don't know. There are vagueposts on X about this model running on two DGX Sparks. If they did the same scale-up with Air instead of Flash, that would probably be about right.
Re: Ox-Alpha Is GLM?
#54If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?
Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5
> 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.Re: Ox-Alpha Is GLM?
#55https://files.catbox.moe/k52n6k.png
The upper chart shows the availability of Ox Alpha and the lower chart shows the availability of GLM 5.3 by Z.ai. They had a blip at exactly the same time.
Re: Ox-Alpha Is GLM?
#56My bet: it's Google running a "new" model based on GLM
Interesting. I like this theory because it explains why some G employees were vague posting about it. But like.. why? Why wouldn’t Google just use Gemma?
They were just trolling.
Re: Ox-Alpha Is GLM?
#57My strong prediction: this model is around 64B and can run on laptops. Thats the reason behind the hype.
We're lucky if it'd fit in 1 DGX Spark. Laptops - nah, unless you mean like an M5 Max with 128GB of RAM then maybe.
> Thats the reason behind the hype.
The hype is imagine DeepSeek Flash before the price increase with even better performance. It'd be like unlimited Sonnet.
Re: Ox-Alpha Is GLM?
#58Re: Ox-Alpha Is GLM?
#59It is probably from Google and is probably hosted on Vertex AI. Opencode announced that responses from Ox Alpha should be better and soon posted about Vertex eu and us multi region update in their changelog. I could be a Gemini model or a new one based on GLM based on tokenizer. Also the amount of inference it is providing for free is something only google can support with its TPUs. So maybe a GLM based model running…
Re: Ox-Alpha Is GLM?
#60Earlier quoted context omitted.
Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5
> 1 quadrillion tokens per day on Nous portal If you are referring to this number ( https://xcancel.com/NousResearch/status/2090899914700054780 ), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.
OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day.
https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300