Live data from Hacker News

GLM-5.3-Flash

z.ai

391–400 of 605 posts

Re: GLM-5.3-Flash

#391
although i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year.

I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas.

In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year.

Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7.

It's not that crazy of an idea and the numbers aren't that bad.

Re: GLM-5.3-Flash

#392
post #186

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want.

This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch...

Edit: yes company money. We don't get subscriptions we pay per token.

Re: GLM-5.3-Flash

#393
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

Qwen 3.8 27B is around Opus 4.8 level of capability on the Agentic Intelligence Index (52 vs 57). In my testing the locally hosted Qwen is good enough that looking at a given piece of work output I couldn't tell you which model was behind it. https://artificialanalysis.ai/models/qwen3-8-27b?models=gpt-...

Qwen3.8 27B (which I adore) is nowhere near Opus 4.8 at puzzle games testing fluid intelligence, https://quesma.com/blog/baba-is-aug-2026/

Re: GLM-5.3-Flash

#394

Earlier quoted context omitted.

If only you could grab those models and host them literally anywhere else where you wouldn't be subject to those terms. Damn. Maybe we'll have to wait for someone to invent something like open download of model weights.

Care to buy me a $10,000-$100,000 computer?

No but a provider with more amicable terms can.

Re: GLM-5.3-Flash

#395

Earlier quoted context omitted.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

Compare to the cost of professional-grade tools in other trades and craft hobbies. Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies. And for some people, $4000 for a device you have complete control…

Yeah, in any other profession where you need to buy a van to drive stuff around, you easily spend similar amount of money on capital investment.

Re: GLM-5.3-Flash

#396

Earlier quoted context omitted.

And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

They are still industry leaders. They'll have to try to maintain that.

Re: GLM-5.3-Flash

#397

Earlier quoted context omitted.

And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…

I was curious about Ox Alpha yesterday, so tried the Tiananmen Sq and got an accurate answer from a third-party player with a little interface on what is claimed to be Ox Alpha: https://oxalpha.com/chat?q=what+happened+in+Tiananmen+square... (and a more detailed answer today when I asked again).

But nothing (at all) from asking GLM-5.3-Flash directly in the OpenRouter chat interface.

Re: GLM-5.3-Flash

#398

Earlier quoted context omitted.

And don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?

I'd like to try some different models, but I've heard that models from China are censored. A government enforced distortion field is a nonstarter for me. To test the waters, I tried the following prompt for each: "What historical event is Tiananmen Square most closely associated with?" Deepseek: I am sorry, I cannot answer that question. I am an AI assistant designed to provide helpful and harmless responses. GLM-5.3…

Deepseek and GLM answered correctly on Openrouter when using non-Chinese endpoints. I hope it stays that way!

Re: GLM-5.3-Flash

#399
post #348
post #186

Earlier quoted context omitted.

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

Opus 5 is the least reliable frontier-class model in the market

In what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.

Re: GLM-5.3-Flash

#400
post #191

Earlier quoted context omitted.

What they don’t do. They claim that the GDPR applies if you provide your service in the EU, and that’s a valid claim.

They obviously don’t have a valid claim to be able to tell everyone who wants to put a website on the internet that they have to do it the EU way, which is what we’re actually talking about.

If the site can be viewed in the EU, it has to follow EU rules. Not different from other countries.

What do you think why the normal polymarket site is blocked for US users.

Post reply on HN