Live data from Hacker News

GLM-5.3-Flash

z.ai

171–180 of 605 posts

Re: GLM-5.3-Flash

#171

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

Chinese laws are not valid in the EU

That’s pretty funny to say when the EU claims GDPR applies worldwide.

Re: GLM-5.3-Flash

#173

Earlier quoted context omitted.

> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of…

I have exactly the same opinion Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled…

These are great points. It's a little off topic but what you bring up is why i advise new grads to spend the first couple years of their career in small eat-what-you-kill companies. I think software devs who start out in large companies get this distorted view that their twice a month direct deposit is just magic and comes from the ether no matter what they do. The whole industry would be better off if everyone started out in a "you don't deliver, you don't eat" company and grew from there.

Re: GLM-5.3-Flash

#174
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini) And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet…

I can easily see a situation where most non American AI usage is on Chinese models on Chinese chips though.

Re: GLM-5.3-Flash

#175

Earlier quoted context omitted.

There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.

It cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.

> A GPU in your basement is a depreciating asset

All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.

Re: GLM-5.3-Flash

#176
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini) And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet…

What about the current situation, where serious API payers are increasingly OK with using open-weight models running on US providers?

https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...

Re: GLM-5.3-Flash

#177
Will we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out?

Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?

Re: GLM-5.3-Flash

#178
post #104

Earlier quoted context omitted.

Well, Luna debuted with 5x higher pricing than is currently available. With the pace of recent development these models might not be relevant by Thanksgiving.

Of course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are…

I don't know if that's the standard pricing for US models to go down overtime, while Chinese ones go up (start cheap but pay more).

I don't have enough metrics to compare those costs but still Chinese models have been cheaper except against Luna for me.

FWIW, Luna does everything so well, I just keep using it for all my agents by default.

Re: GLM-5.3-Flash

#179

Earlier quoted context omitted.

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

Compared to pricing from 3 years ago, it's insane.

The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.

Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.

(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)

Re: GLM-5.3-Flash

#180

Earlier quoted context omitted.

Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

Has there been any confirmation about what that model even is? Edit: Ah: > This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

It's also in this very announcement, in the first paragraph:

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

Post reply on HN