Live data from Hacker News

GLM-5.3-Flash

z.ai

551–560 of 605 posts

Re: GLM-5.3-Flash

#551
post #223

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips. Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified. whether it's t…

Weren't they giving free access? Not exacty a meaningful heuristic if so

The key insight here is being the top used model on opencode while being fully served on Chinese chips. The free price itself might be just a flex or marketing budget.

Re: GLM-5.3-Flash

#552

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?

Re: GLM-5.3-Flash

#553

Earlier quoted context omitted.

Slightly more expensive than the (post-price hike) DS4 flash pricing, but in the ballpark. https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...

Hm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off. EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.

I mean, it implies that it has an even better cache hit rate that DS flash, which is impressive, as the chr on DS flash was already really good in my experience

Re: GLM-5.3-Flash

#554

Earlier quoted context omitted.

What is actually getting you flagged by the openweights inference providers? Thus far I haven't hit any of the reverse engineering or infosec guardrails that Anthropic is so keen on

While I'm sure some of the open weight providers do this as well, I think the comparison is frontier labs v local inference.

I'm not sure that is the comparison - the OP is planning to run an open weights model locally, the obvious comparison would be paying a hosted provider to run the exact same model

Re: GLM-5.3-Flash

#555
post #518

Holy shit, is this model really that bad??? Just asked it a question via the custom opencode go endpoint routed over cloudflare ai gateway doesnt show me the correct models. My fault was that I set " https://opencode.ai/zen/go " as endpoint and tried my-gateway.com/custom-ocgo/v1/models, turns out I had to add /v1 to the opencode url and leave it on my -gateway.com. But first it told me that opencode is not on the co…

I actually noticed this too with DS flash over openrouter. Instructions to respond in my native language cause the model to spew gibberish roughly resembling my native language mixed up with similar languages, and it was also dropping in Cyrillic characters randomly (although i noticed this with the anthropic models as well)

Re: GLM-5.3-Flash

#556
post #301

Earlier quoted context omitted.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

They only went up from 3000€ to 4000€ which isn't a lot. For comparison the cheapest Strix Halo 128GB went from 1600€ to 2600€ in the same timeframe.

> They only went up from 3000€ to 4000€ which isn't a lot.

Keep in mind that the DGX Spark was delayed quite a while, meaning that starting 3k price tag is already well into the RAM crisis - just 2 years ago, a 128GB DDR5 kit could be had for $600

Re: GLM-5.3-Flash

#557

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

Anthropic has banned me for using Claude via a VPN.

They don't even allow me to download my data.

How is it better?

Re: GLM-5.3-Flash

#558

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

I blocked Z.ai as soon as they were loading 10 different external providers including Alibaba who was just proven to execute silent sound fingerprinting mechanisms.

Does it work with NoScript?

Re: GLM-5.3-Flash

#559
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.- Further quote: "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference effic…

> comparable to mainstream NVIDIA GPUs

By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models

Re: GLM-5.3-Flash

#560
post #426
post #228

Earlier quoted context omitted.

I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…

I can't get 3.8 to exit thinking loops. It will just think and think and think on the most trivial topics. I wanted it to port a speed test powershell script to c#. Claude opus 5 completes it under 60 seconds. I let 3.8 churn about 6 different times for 30+ minutes and it never wrote a single line of code to disk. It wrote lots of lines in thinking. unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL DSH (pi) Any tips?

Use Muse Glimmer. It’s good.
Post reply on HN