Live data from Hacker News

GLM-5.3-Flash

z.ai

571–580 of 605 posts

Re: GLM-5.3-Flash

#571

Earlier quoted context omitted.

Casual consumers are using American models because their usage is low. As usage scales, the economics heavily favor open weight models. The API pricing from American companies is absurd. This is particularly true in an enterprise setting.

Open weight model hosts don't have the compute to meet enterprise demand. A large part of why these models are so cheap is because overall demand for them is incredibly low. Back in May, Gemini alone was doing about a month's worth of Openrouter tokens every day.

I disagree totally. DeepSeek raised prices because they couldn’t serve the demand. But there are tons of American vendors ready to fulfill it. Many enterprises, including the one I work for, are swapping to open weights.

Why wouldn’t you?

Re: GLM-5.3-Flash

#572
I guess this is also a really good indicator of just how much more tokens people would use when not limited financially. Wonder how much that tells the labs about pricing. Then again, it's 2.3x not 10x the next paid/cheap model.

Re: GLM-5.3-Flash

#573
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology. Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium t…

Yeah, I just ignore the benchmarks at this point. For open-weight models the provider's setup impacts performance so you can have different experience's with the same model at the same quantization from different provider's. Just have to use them on real tasks with your actual harness to really know how they will perform and hope the provider doesn't do something to degrade performance (e.g. update the middleware to a new version with a defect that impairs performance).

Re: GLM-5.3-Flash

#574

Earlier quoted context omitted.

I have exactly the same opinion Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled…

This is really insightful, thank you. Can you share a bit more about how you shifted to be more aligned with P/L? And how to accurately estimate incremental upside? I'm an early PhD student with interest in ibdustrial research/R&D, and currently struggling to understand how to think about how to navigate through my career.

[deleted]

Re: GLM-5.3-Flash

#575

Earlier quoted context omitted.

> It's too big, bright and resourceful of a country to choose confrontation instead of collaboration. It's not like we didn't try it. China first have to learn to make deals where both party benefits.

I would say the Trump Administration needs to learn this as well.

America knew how to do it. But they are learning quick from China.

Re: GLM-5.3-Flash

#576

Earlier quoted context omitted.

I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...

DS4 Flash absolutely kicks ass for reverse engineering and bug hunting. Almost no point in considering paying for a bigger model, although it's possible the stuff I've fed it (wide variety of older DOS/Windows stuff and device firmwares) might be easier targets.

idk if you'll read this, but can you explain more the setup needed for RE?

Re: GLM-5.3-Flash

#577

Earlier quoted context omitted.

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…

I ran this against the version of deepseek v4 flash 0731 running in fireworks ai and it responded with this: Tiananmen Square has been the site of many major historical events, but internationally it is most closely associated with the 1989 pro-democracy protests and the Chinese government's military crackdown on protesters there in June of that year, which resulted in many deaths and injuries. The event remains a se…

That's encouraging, thank you!

Re: GLM-5.3-Flash

#578

Earlier quoted context omitted.

I've been reverse engineering LEON3-FT SPARC v8 BE code, so I wouldn't say it's common :D. When attached to Ghidra through a MCP the things you can do with this are simply crazy.

https://github.com/bethington/ghidra-mcp . Works flawlessly.

https://github.com/bethington/ghidra-mcp/issues/307

Re: GLM-5.3-Flash

#579
post #393

Earlier quoted context omitted.

Qwen3.8 27B (which I adore) is nowhere near Opus 4.8 at puzzle games testing fluid intelligence, https://quesma.com/blog/baba-is-aug-2026/

yeah it's more like opus 4.6 iirc

Not really compatible on all fronts, it's very capable especially with tool calling, workflows, logic and its base coding ability, but it's only a 27b model so it does not have anywhere near the level of knowledge baked in as larger models. This does not mean that it's not a good or useful model - it is on both accounts and very efficient, but it's not similar to a large model generally speaking.

Re: GLM-5.3-Flash

#580

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training…

They innovated a lot.
Post reply on HN