Live data from Hacker News

GLM-5.3-Flash

z.ai

81–90 of 605 posts

Re: GLM-5.3-Flash

#81
post #62
post #58

Earlier quoted context omitted.

Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time. I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

seems unlikely that they'll get nearly as much demand now that it isnt free

Sure, although I still expect it to become the most used model on openrouter.

Re: GLM-5.3-Flash

#82
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

It's also better than Sol (at whatever effort) at designing pretty UIs. I have a Codex sub and I've been using this model for UI stuff.

Re: GLM-5.3-Flash

#83
post #38
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

Another self-inflicted own courtesy of US government policy. While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.

That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.

Re: GLM-5.3-Flash

#84

Earlier quoted context omitted.

If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?

There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.

For me it's entirely because I have a bunch of projects with my own personal data that would be tough to do with openrouter/claude or any other cloud.

For example, I have a small posix-shell-based LLM harness that can SSH into my NAS and run organization tasks using the local DS4Flash that I have right now. It's already been a massive help for me to keep me organized, and that's just 2x DGX Spark's worth of compute.

Re: GLM-5.3-Flash

#87

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute). I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends…

We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.

Re: GLM-5.3-Flash

#88
post #46

Earlier quoted context omitted.

I thought 4000 in sum. No wait, 4000 per , plus tax. Or EUR pricing to similar accord. Ouch.

Yeah, that little cluster costs about the same as a brand-new Dacia Sandero.

to save others from having to look up what that is like I did, it's a car

Re: GLM-5.3-Flash

#89

Earlier quoted context omitted.

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

I mean, I have the same machine and the pricing is only what it is because it has that 1TB nVME in it instead of larger. nVME prices are insane and have been for months. Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.

I already have a recycled QNAP-now-TrueNAS that has 10G connections into the fleet so the 1TB doesn't bother me at all. I did some rough math and I don't think I'll ever need to load weights off NFS for what I'm doing so far, but the capability is there.

Re: GLM-5.3-Flash

#90
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)

And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.

So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.

I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.

If anything it's custom chips from the labs that threatens Nvidia.

Post reply on HN