Earlier quoted context omitted.
Is that cheaper than DS4 flash?
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
GLM-5.3-Flash
21–30 of 605 posts
Re: GLM-5.3-Flash
#22> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. From a biased source, but would be big if true. I've had great results with GLM 5.2. From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
Re: GLM-5.3-Flash
#23Re: GLM-5.3-Flash
#24> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
>> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."
It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.
Re: GLM-5.3-Flash
#25I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field. It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
Re: GLM-5.3-Flash
#26Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?
Re: GLM-5.3-Flash
#27Re: GLM-5.3-Flash
#28For those who didn't read, this is the identity of the mysterious "Ox Alpha" model
│ https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash.
│ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}
Re: GLM-5.3-Flash
#29> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.
It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.