Live data from Hacker News

GLM-5.3-Flash

z.ai

561–570 of 605 posts

Re: GLM-5.3-Flash

#561
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology. Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium t…

For me it's been this: Opus (at least in my experience) is unbeatable at "planning the work". That includes a lot of things, including getting arch. sorted, a chassis/skeleton done. Filling that up and doing the actual "coding," though, I've noticed no real difference between the Claude model and GLM. So it'll be interesting to see how this flash model compares cost-wise to what I'm currently using, which is GLM 5.3, for coding. Looks like it will be reduced even further and might be great if it's faster (and better?) than 5.3 in my real world/personal experience.

(I'm someone who doesn't really care about delays of a few seconds, or even more than few seconds. But if I am trying to notice then sure Claude is definitely faster as well).

Re: GLM-5.3-Flash

#563

Earlier quoted context omitted.

Hm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off. EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.

I mean, it implies that it has an even better cache hit rate that DS flash, which is impressive, as the chr on DS flash was already really good in my experience

I think the weighted average takes into account all providers (some DS4 flash providers are 'premium' providers and offering higher speeds for higher pricing) and these are tilting the scale

Re: GLM-5.3-Flash

#564

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training…

Presumably there is distillation or similar being used to transfer from a larger model to a smaller one.

Re: GLM-5.3-Flash

#565
GLM-5.3-Flash is now the cheapest way to reach Artificial Analysis Intelligence Index 52.3 to 57.5, at 8.7¢ per task, and displaced Muse Spark 1.2, Grok 4.5, Gemini 3.7 Flash, and DeepSeek V4 Pro from the Pareto frontier.

advance card: https://catalystneuro.com/llm-cost-frontier/images/advances/...

tracker: https://catalystneuro.com/llm-cost-frontier/

Re: GLM-5.3-Flash

#567
post #248
post #185

Earlier quoted context omitted.

I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…

The one time I tried asking Sol to use subagents for a small project, it took a surprisingly long time, used up the entire usage limit in one go, and basically failed the project. I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickl…

OpenAI has been doing wonky stuff with subagents, including encrypting the prompts sent to subagents in Codex. Who knows what’s really going on.

Re: GLM-5.3-Flash

#569

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

This is a China derived model. Interactions with different leadership styles, governments, datasets, and culture impact the resulting AI.

I suspect any model connected or built by China will serve their interest despite their terms and agreements.

Re: GLM-5.3-Flash

#570

I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field. It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

> It's too big, bright and resourceful of a country to choose confrontation instead of collaboration. It's not like we didn't try it. China first have to learn to make deals where both party benefits.

I would say the Trump Administration needs to learn this as well.
Post reply on HN