Live data from Hacker News

GLM-5.3-Flash

z.ai

331–340 of 605 posts

Re: GLM-5.3-Flash

#331
post #186

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

You'll find that hard to prove objectively and conclusively.

Re: GLM-5.3-Flash

#332

So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.

On Twitter they mentioned that it was unfortunate timing as the 3.7 flash release collided with ox alpha.

Re: GLM-5.3-Flash

#333
i wonder if more companies will now stealth launch their models. imagine they just released this on openrouter for free but under their normal name - would they get the records in token usage then?

Re: GLM-5.3-Flash

#334

So the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.

Trolling. GLM is heavily distilled from Gemini.

Early glm models gave off that wibe. But now its more inspired by. With a bit of Claude in there. But I do think they actually do RL otherwise glm5.3 shouldn't have been able to beat fable on the few tests it did.

Re: GLM-5.3-Flash

#335
Does the word "flash" mean a specific thing when it comes to LLM models? I noticed that this word is used by gemini, qwen, and z.ai and I'm curious does it mean the same thing for each one, or did they all just accidentally brand similarly?

Re: GLM-5.3-Flash

#336

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

How is that vastly different from any other non-enterprise facing provider? I do believe Anthropic bans accounts without even a human in the loop with no recourse left to those banned.

Re: GLM-5.3-Flash

#338

Earlier quoted context omitted.

This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.- Further quote: "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference effic…

Anyone knows what are those Chinese chips? Can they be bought? (Assuming im not i the US, And actually im in a 3rd world country).

They are high end really expensive Huawei ascend GPUs. It is kinda bruteforcing the performance on a older semiconductor processing tech, so total production is pretty low.

Re: GLM-5.3-Flash

#339
post #335

Does the word "flash" mean a specific thing when it comes to LLM models? I noticed that this word is used by gemini, qwen, and z.ai and I'm curious does it mean the same thing for each one, or did they all just accidentally brand similarly?

It seems like it has come to mean "fast, small, cheap" models these days, and seems well enough understood as such that different AI labs are adopting it.

Re: GLM-5.3-Flash

#340

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures

I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

Post reply on HN