Live data from Hacker News

GLM-5.3-Flash

z.ai

21–30 of 605 posts

Re: GLM-5.3-Flash

#21
post #8

Earlier quoted context omitted.

Is that cheaper than DS4 flash?

All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.

Tbh it was also slow because it was being hammered by everyone making use of the free tokens

Re: GLM-5.3-Flash

#22
post #16

> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. From a biased source, but would be big if true. I've had great results with GLM 5.2. From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

> From a biased source, but would be big if true. I've had great results with GLM 5.2.

It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.

Re: GLM-5.3-Flash

#24
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

This is no surprise [0] [1].

>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."

It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.

[0] https://news.ycombinator.com/item?id=49397204

[1] https://news.ycombinator.com/item?id=49431231

Re: GLM-5.3-Flash

#25

I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field. It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.

Re: GLM-5.3-Flash

#26
> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

Re: GLM-5.3-Flash

#28

For those who didn't read, this is the identity of the mysterious "Ox Alpha" model

They even give this over the API now:

https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.

│ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}

Re: GLM-5.3-Flash

#29
> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

Post reply on HN