Live data from Hacker News

GLM-5.3-Flash

z.ai

201–210 of 605 posts

Re: GLM-5.3-Flash

#201

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute). I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends…

> I don't see how NVIDIA can keep their spot as belle of the ball.

FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications.

With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.

Re: GLM-5.3-Flash

#202

Earlier quoted context omitted.

> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of…

To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.

Use Opus 4.8. 5 is absolute garbage.

Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.

Re: GLM-5.3-Flash

#203
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.

Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)

Re: GLM-5.3-Flash

#204
post #82

Earlier quoted context omitted.

It's also better than Sol (at whatever effort) at designing pretty UIs. I have a Codex sub and I've been using this model for UI stuff.

> I've been using this model for UI stuff. The flash one?

Yes. I have no UI experience, and wanted a model that could produce something good without me telling it how anything should look like.

My prompt was something like: "here's data I have, here's what matters to me, create HTML mockup".

All GPT 5.6 models were laughably bad. And I don't want to downplay it - they were just absolutely, objectively horrible. Every single attempt was what I could probably call "if json was ui".

Claude models produced... "claude look".

GLM 5.3 - somewhere between GPT and Claude.

Kimi k3 - each attempt produced beautiful UIs. It used components that I didn't even know existed and wouldn't even know to ask for. But expensive, very expensive.

ox-alpha (GLM 5.3 flash) was very close to K3. And at this price point, it's already configured as "designer" model in my oh-my-pi.

Re: GLM-5.3-Flash

#205

Earlier quoted context omitted.

Of course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are…

I don't know if that's the standard pricing for US models to go down overtime, while Chinese ones go up (start cheap but pay more). I don't have enough metrics to compare those costs but still Chinese models have been cheaper except against Luna for me. FWIW, Luna does everything so well, I just keep using it for all my agents by default.

I haven't noticed the Chinese models going up in price for the same model. They do release new versions of the models with different prices that are higher. But everybody is doing that. One fine point is that deepseek-v4-flash-0731 is really a different model than deepseek-v4-flash and it's priced higher.

Re: GLM-5.3-Flash

#206
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?

If your usage wouldn't change with local inference and you don't have security/privacy concerns then at the currently heavily subsidized pricing, sure.. not economical.

But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".

I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.

Re: GLM-5.3-Flash

#207

Earlier quoted context omitted.

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

> $4,000 isn't priced insanely? ye gads

It depends.

My bicycle was in the 5-digits brand new (now I paid it 1/5th of that and I do thank the first owner for that: the 8 000 out of 10 K I saved were put into stocks, that's his opportunity cost, not mine).

Or I know a great many a going to cry "audiofool", but I can say with certainty the following does sound better than the stereo setup of those crying audiofool:

https://youtu.be/TQg9FTBMcTQ

(not my setup but I've got those speakers: same thing, 15 K EUR brand new for the pair... Previous owner forked the money to buy these brand new and, well, I didn't... And I just hooked them to a wonderful, cheap, fully-integrated Yamaha amp: amazing sound).

If your hobby is DIY job around the house, the cost of tools can very quickly add up too: having 20 K worth of tools is definitely not unthinkable.

You like old cars? Pricey hobby.

Some here even track their cars: tires and brake pads budget (and overall car budget and depreciation)... Through the roof.

There's a saying that you're not really into computers if your setup doesn't cost more than your car.

Is $16 K ($4 K x 4) a lot? It's six months of rent for me and for many here I'm sure. It's not "crazy crazy".

Can anyone afford that? Definitely not. But there are way more insane things out there.

And thanks to the individuals that go through to all the pain of setting those up, we've got feedback, tutorials, explanation, numbers, etc. as to how to run those at home.

For example I helped my brother set up VMs and GPU passthrough and he's now running uncensored models locally and showing me the different answers between the uncensored models and the commercial, censored, ones.

So to GP who bought four of these: we need more people like you on HN, keep it going, blog about it, be "crazy"!

Re: GLM-5.3-Flash

#208
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

Hopefully you also bought a switch

they have 2 interfaces each so you typically daisy chain them

Re: GLM-5.3-Flash

#210

Earlier quoted context omitted.

There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.

It cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.

At the current point in time I'd argue it's more about opportunity cost/value.

If I'm a professional photographer chasing the best possible end product, I'm not buying cameras because they're economical. I'm buying the best camera I can get my hands on to get the best product I can produce within reason under the understanding that it doesn't have to equate to the best economic decision to be the _right_ decision.

If you're in a position to be able to take advantage of the local inference - it's a no brainer. If you're not sure how that would be done, then it's not a good move.

Post reply on HN