Live data from Hacker News

GLM-5.3-Flash

z.ai

431–440 of 605 posts

Re: GLM-5.3-Flash

#431
post #100

Earlier quoted context omitted.

Isn't this practically every TOS though? Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it. HN's for example > We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time. > You acknowledge that Y Combinator may establish general practices and limits concerning use o…

> Isn't this practically every TOS though? Not even close. Even OpenAI and Anthropic aren't bad enough that they claim literal ownership of your inputs and outputs. > HN's for example You're not paying to use HN. Getting banned here has essentially zero consequences. If Z.ai uses its absolute powers to ban you because you wrote a review about them or something, then you lose actual money. This is especially relevant…

> You're not paying to use HN.

Oh for sure man, this absolutely looks like you were only concerned and talking about paid services:

> Broad and perpetual license over inputs and outputs, and even your name and profile picture.

Re: GLM-5.3-Flash

#432
post #120
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

> They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts It's what people know. Opus is just the common target. > Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash The problem with this and DeepSWE is it goes for a very specific profile. I'm not conv…

> I'm not convinced DeepSWE is any accurate in actual work.

They listed Muse Spark 1.2 around DeepSeek V4 Flash even though it's a much shittier model in basically every aspect.

> GLM is a better all rounder in some ways. Better at creativity.

I agree with the creativity part.

Re: GLM-5.3-Flash

#433
tbh I wasnt that impressed by it. initial benchmarks were trying to say it was AGI but i told it to re-build Palantir in 1 pass and it gave me a non working prototype

Re: GLM-5.3-Flash

#434

Earlier quoted context omitted.

Compare to the cost of professional-grade tools in other trades and craft hobbies. Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies. And for some people, $4000 for a device you have complete control…

That's only half the reason it's expensive. The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.

> The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider

That's just a one-dimensional thought! Your own hardware gives you complete control, and it doesn't time you out for 4 hours, unlike those vendors.

Re: GLM-5.3-Flash

#435
post #185
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…

I only use Luna (max), I find it very rarely just reasons. In fact, I find it reasons too little.

Re: GLM-5.3-Flash

#436
post #5

> with all of this traffic served on Chinese AI chips RIP Nivida shareholders

Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

Ox Alpha was also serving 10T+ tokens a day for free.

When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

Re: GLM-5.3-Flash

#437
post #87

Earlier quoted context omitted.

We need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildo…

Has nobody from any of the companies hosting open weights models released detailed information on how much it really costs?

I’m sure someone does, but I’ve never seen anything other than vague statements like Anthropic’s “inference is profitable” comment. I suspect everyone is playing everything close to the vest because they aren’t yet public and they want to control the information flow to the street.

Re: GLM-5.3-Flash

#438
post #186

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

Honestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.

Re: GLM-5.3-Flash

#439

Earlier quoted context omitted.

Wow that's uncanny.

Given it was probably one of the simplest things you could change in the codebase, the kind of stuff you give a new developer on the project, I'm not sure it's so telling, there is usually just about one way to remove a feature flag.

The specs that were written are what I found surprising, not that it removed the flag in the same way.

Re: GLM-5.3-Flash

#440

This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

There is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.
Post reply on HN