Live data from Hacker News

GLM-5.3-Flash

z.ai

611–612 of 612 posts

Re: GLM-5.3-Flash

#611

Earlier quoted context omitted.

> DeepSeek token prices are continuing to _increase_ One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...

You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.

[flagged]

Re: GLM-5.3-Flash

#612

Earlier quoted context omitted.

I've found it tends toward long thinking loops even for simple tasks (and any quantization seems to increase their length), but those do exit eventually, unlike with Qwen 3.6. I use the Unsloth UD_Q2_K_XL GGUF with default parameters, along with that custom template linked elsewhere in the thread, and no K/V cache quantization.

For smaller models, you'll probably find anything below Q4 will need handholding. Check Unsloth's graphs at the different quantisations VS error rates and you'll see why.

Thank you, I know about those. And I'll stick with this quant. Normally I'd be with yout there, but Qwen 3.8 is turning out to be good at self-correcting, and the free memory I can use for extra context is worthwhile.
Post reply on HN