> 320B total parameters and just 18B active parameters This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home. @edit: so many releases that I forgot to math. This fits j…
GLM-5.3-Flash
31–40 of 605 posts
Re: GLM-5.3-Flash
#32Re: GLM-5.3-Flash
#33By clicking this link you download some PDF in the background
Re: GLM-5.3-Flash
#34> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
Re: GLM-5.3-Flash
#35It's only 320B, local frontier AI is getting closer, sooner than expected.
Re: GLM-5.3-Flash
#36> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. From a biased source, but would be big if true. I've had great results with GLM 5.2. From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
Re: GLM-5.3-Flash
#37Re: GLM-5.3-Flash
#38> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.
Re: GLM-5.3-Flash
#39I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field. It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
Re: GLM-5.3-Flash
#40Earlier quoted context omitted.
All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
Tbh it was also slow because it was being hammered by everyone making use of the free tokens