> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
GLM-5.3-Flash
11–20 of 605 posts
Re: GLM-5.3-Flash
#12Re: GLM-5.3-Flash
#13Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…
Wow, if you don't mind me asking. How and where?
Re: GLM-5.3-Flash
#14> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
Further quote:
"Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."
Re: GLM-5.3-Flash
#15Re: GLM-5.3-Flash
#16From a biased source, but would be big if true. I've had great results with GLM 5.2.
From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
Re: GLM-5.3-Flash
#17Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…
> get myself four sparks at a decent price Wow, if you don't mind me asking. How and where?
They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
Re: GLM-5.3-Flash
#18> 320B total parameters and just 18B active parameters This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home. @edit: so many releases that I forgot to math. This fits j…
Re: GLM-5.3-Flash
#19Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
Is that cheaper than DS4 flash?