Live data from Hacker News

GLM-5.3-Flash

z.ai

61–70 of 605 posts

Re: GLM-5.3-Flash

#61
post #38

Earlier quoted context omitted.

Another self-inflicted own courtesy of US government policy. While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

The export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah

> The export controls were revoked before

Zai is on another "export control" list outside the broader 1. Doesn't help.

Re: GLM-5.3-Flash

#62
post #58
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time. I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

seems unlikely that they'll get nearly as much demand now that it isnt free

Re: GLM-5.3-Flash

#63
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.

Re: GLM-5.3-Flash

#65
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

Re: GLM-5.3-Flash

#66

Earlier quoted context omitted.

Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.

That's how you catch up when you're behind. Now the US is behind in EVs can you guess what they're doing? [1] [1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...

"argument is very weak" regardless as I said.

Re: GLM-5.3-Flash

#67
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of…

To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.

Re: GLM-5.3-Flash

#68
post #8
post #3

Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03

Is that cheaper than DS4 flash?

Slightly more expensive than the (post-price hike) DS4 flash pricing, but in the ballpark.

https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...

Re: GLM-5.3-Flash

#69

Earlier quoted context omitted.

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

Compare to the cost of professional-grade tools in other trades and craft hobbies.

Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.

And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.

Re: GLM-5.3-Flash

#70
post #8

Earlier quoted context omitted.

Is that cheaper than DS4 flash?

It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now

Few weeks ago, I wouldn't expect this statement to be true. Accelerate!
Post reply on HN