Live data from Hacker News

GLM-5.3-Flash

z.ai

441–450 of 605 posts

Re: GLM-5.3-Flash

#441
post #186

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

Eh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper.

The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.

Re: GLM-5.3-Flash

#442

Earlier quoted context omitted.

> Isn't this practically every TOS though? Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs. > HN's for example You're not paying to use HN. Getting banned here has essentially zero consequences. If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant…

> You're not paying to use HN. Oh for sure man, this absolutely looks like you were only concerned and talking about paid services: > Broad and perpetual license over inputs and outputs, and even your name and profile picture.

The post I was replying to wasn't talking about that, but sure, let's consider it.

Everything I post here is public, and it's just relatively low value commentary anyway. It doesn't matter if Y Combinator has rights to it. Arguably they actually need to assert some rights, otherwise they wouldn't be able to transfer copies of copyrighted comments to other visitors of the site. HN's terms are probably too broad for their purposes but it doesn't really matter much because this is just a forum.

AI on the other hand is for actual work, both public and private. There are actual economic implications here, so the stakes are much higher. I absolutely want to own the inputs and the outputs I paid money for.

Re: GLM-5.3-Flash

#443

Earlier quoted context omitted.

You aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify). You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most. Then there…

yeah i mostly agree, especially compared to subsidized subscription cost. But for a heavy user who has enough work to be done so that the box runs almost 24/7 at say 50tok/sec, the math gets interesting against API prices. And it can be interesting compared to subscription in the sense that you don't have the quota anymore. That means there's probably a lot of things you're not doing because of the quotas that you co…

At 50tps for single stream you are going to get 50 * 60 * 60 * 24 * 30 = 130M out tokens of GLM 5.3 Flash...

That's less than what 40$ at current API rates... So if you are willing to pay 200$ per month you will get much better limits paying API rates.

You can't run large Kimi K3 models on 10K worth of hardware either way, you need to spend like 50K USD minimum.

Just pay for the API rates or get a low cost provider that uses higher batching, you can get shittier tps but much better prices, probably go as low as 20$ for as much usage as you can ever get from a 10K USD machine from GLM 5.3 Flash...

The issue is nothing expensive runs on these devices and cheap stuff isn't worth running locally, eletricity costs ~12cents/kwh in us iirc, so at 330W M5 Ultra will burn around 8 * 0.12 = ~1$ per day extra in electricity so the electricity is going to cost you the same as the API rates(30$ per month).

I truly don't think you are accounting for the costs here properly. But again if money truly doesn't matter it's much better for privacy and better than paying one of the shady AI labs who are doing god knows what with your data.

Re: GLM-5.3-Flash

#444

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

[deleted]

Re: GLM-5.3-Flash

#445
post #228

Earlier quoted context omitted.

As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely useless at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant st…

I've had the exact opposite experience. I've been using 3.8 for my daily driver since last week, and I've gradually been giving it more and more complex tasks as it continues to deliver high quality results. Now I am basically handing off large complex features, and 3.8 is doing the planning, task breakdown, implementation and review with just a few notes from my side. The tradeoff is time (especially on RDMA4 hardwa…

Both of you should mention what quant you're using. And as another comment said, what tasks you're doing, i.e. coding, classification, summarizing etc.

Re: GLM-5.3-Flash

#446
post #419

Earlier quoted context omitted.

Yes Chinese models censor some historical events. This is nothing knew and well known thing. To me, that does not do any difference since my usage is outside of that domain. Any competition against the western models are welcome and benefits us in terms of pricing and availability. If they have to comply with CCP to be able to do it, then so be it. I have zero sympathy for Anthropic and OAI being so secretive and act…

What it shows is that the CCP has enough oversight and control (either explicitly or by the companies making these decisions by default) that they will alter the models to benefit China. Who is to say they aren't doing it in other ways as well? That they aren't, or won't be, subtly hamstrung in engineering work? OAI and Anthropic have their own issues, you're right to be suspicious of them, but it's not like their mo…

If you don't want to use the Chinese model, then don't. Why attack it instead? Don't you want others to use it either?

Re: GLM-5.3-Flash

#447

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

There is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.

Dario has less space than a Nomad!

Re: GLM-5.3-Flash

#448
post #417

Earlier quoted context omitted.

The next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures I'm sure they have nothing to rival this on a price/performance basis and have already given up on that

> I'm sure they have nothing to rival this on a price/performance basis and have already given up on that How can you be sure about this? They have unbelievable capital. OpenAI is starting to preview its own chips, which could dramatically change the price/performance. We don't know what else Anthropic has cooked up right now that could rival this if they wanted to. Yes, others will _also_ continue to innovate, but m…

> They have unbelievable capital.

all those Chinese labs are backed by the Chinese government which can just print money.

time to wake up.

Re: GLM-5.3-Flash

#449
post #400

Earlier quoted context omitted.

They obviously don’t have a valid claim to be able to tell everyone who wants to put a website on the internet that they have to do it the EU way, which is what we’re actually talking about.

If the site can be viewed in the EU, it has to follow EU rules. Not different from other countries. What do you think why the normal polymarket site is blocked for US users.

I’ve got a bunch of sites that can be viewed from (checks notes) the internet. If people in Europe don’t like that, they can block it, or choose not to visit it. In the meantime, you’ve proven exactly what I first said, which is that they are claiming it applies worldwide. Here I am, just putting a site on the internet, and you’re telling me I have to follow EU laws concerning it. Nope. I don’t.

Re: GLM-5.3-Flash

#450
post #417

Earlier quoted context omitted.

> I'm sure they have nothing to rival this on a price/performance basis and have already given up on that How can you be sure about this? They have unbelievable capital. OpenAI is starting to preview its own chips, which could dramatically change the price/performance. We don't know what else Anthropic has cooked up right now that could rival this if they wanted to. Yes, others will _also_ continue to innovate, but m…

w.r.t. the OAI chips, wouldn't they be subject to the same bottlenecks that has plagued semis lately or at least be forced to pay a pretty premium to circumvent that?

They'll be able to buy them without paying NVidia's 80% profit margin
Post reply on HN