Live data from Hacker News

GLM-5.3-Flash

z.ai

321–330 of 605 posts

Re: GLM-5.3-Flash

#321

Will we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out? Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute deser…

If it gets more efficient it'll be more enticing to expand use case. Personally I'm hoping to do a lot at home but I'm not counting the datacenter building as a bad move at this moment. It may and up that way.

Re: GLM-5.3-Flash

#322

Earlier quoted context omitted.

Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market. Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)

> Most US companies that have anything to do with government, finance, medical, etc... That's a huge market. Compared to the rest of the world?

I don't have any pie charts in front of me, but yes, I would estimate it's a decently big slice of the world market.

Re: GLM-5.3-Flash

#323

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

You can download the weights, and run in your own hardware and avoid all that.

Re: GLM-5.3-Flash

#324

Earlier quoted context omitted.

If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?

There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.

I'm not trying to say there is no use case. I just want to know the cost. Is it less than the API cost? Is it the same? Is it more? I'm looking for hard numbers. If the cost is the same or more, then the decision for local isn't to save money

Re: GLM-5.3-Flash

#325
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

How much did you drop on these 4 sparks?

Sparks + cables + 10g SFP+ came out to ~$21,500 CAD

Re: GLM-5.3-Flash

#326
post #191

Earlier quoted context omitted.

That’s pretty funny to say when the EU claims GDPR applies worldwide.

What they don’t do. They claim that the GDPR applies if you provide your service in the EU, and that’s a valid claim.

They obviously don’t have a valid claim to be able to tell everyone who wants to put a website on the internet that they have to do it the EU way, which is what we’re actually talking about.

Re: GLM-5.3-Flash

#327

Earlier quoted context omitted.

Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology. Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium t…

I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better.

I have had the opposite where I felt 5.3 as a stronger model than 5.2. Its feels way more in tune with my code, and does more nuanced edits. Though I do handhold my models a lot, so might fall outside the agentic term.

Re: GLM-5.3-Flash

#328

Earlier quoted context omitted.

$40,000 GPU is like few pennies in sand. Only mildly hyperbolic. But a GPU fresh out of fab is $2000 after ASML, TSMC and inputs get their 50-75% margin, then somehow $40k laundered through US financialization / Nvidia margins. Commoditized GPUs shouldn't cost more than 1-2% current price once there's competition.

[flagged]

Compelling argument from only msg on new account.

Re: GLM-5.3-Flash

#329
post #141

With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different? Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.

I agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have. For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.

Could do a "Big model for architecture and planning and smaller model (or local model) for implementation" sort of thing

Re: GLM-5.3-Flash

#330

You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…

Yes, the terms are dubious. But they are also reasonably lenient with enforcement. They also don't require persona id verification, witch is wat turned me away from openai.
Post reply on HN