Live data from Hacker News

GLM-5.3-Flash

z.ai

241–250 of 605 posts

Re: GLM-5.3-Flash

#241

Earlier quoted context omitted.

> get myself four sparks at a decent price Wow, if you don't mind me asking. How and where?

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

because it has 1T ssd not 4T

Re: GLM-5.3-Flash

#242
post #102

Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is. What irks me about this is that the harnesses seem to be just an afterthought here. Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, e…

You can use OpenRouter directly in Claude Code as well, it's quite nice!

Re: GLM-5.3-Flash

#243
post #96

Earlier quoted context omitted.

I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plan…

This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does. Not…

> You can't do that on larger models unless you have the hardware like this person does.

Or just rent something substantial for like $4/hr on runpod or w/e to do that.

My gripe is this persons compute is wasteful and makes it harder for me to buy something with like 64gb ram to do normal work and run containers while I keep using cloud models.

Someone else calculated the break even being 10 years, it’s just dumb. And I think it’s clear there won’t be a big rug pull anymore, there are too many open models and providers now.

Re: GLM-5.3-Flash

#244
post #7

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…

I am surprised. I've been using DS4 Flash (0731) for weeks now and it works perfectly fine as a replacement for Claude in a large variety of cases. It requires a few more iterations, sure, but it's useful enough to not need a Claude subscription anymore. Among the things I do I've been reverse engineering, writing complex C++ code...

Re: GLM-5.3-Flash

#245

Earlier quoted context omitted.

> DeepSeek token prices are continuing to _increase_ One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...

You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.

$40,000 GPU is like few pennies in sand. Only mildly hyperbolic. But a GPU fresh out of fab is $2000 after ASML, TSMC and inputs get their 50-75% margin, then somehow $40k laundered through US financialization / Nvidia margins. Commoditized GPUs shouldn't cost more than 1-2% current price once there's competition.

Re: GLM-5.3-Flash

#246

On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M How is the business model of Anthropic/OpenAI will sustain?

This is a bad model. Worse than Luna in every way; slower, dumber.

Re: GLM-5.3-Flash

#247

Earlier quoted context omitted.

I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads

"For new hardware in 2026 with 128Gi of high-speed memory"

Checked a couple days ago and looks like we're at about 3.5x 2020 memory prices (looking at just $/GB).

Re: GLM-5.3-Flash

#248
post #185
post #79

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…

I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…

The one time I tried asking Sol to use subagents for a small project, it took a surprisingly long time, used up the entire usage limit in one go, and basically failed the project.

I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickly and cheaply.

(Is it really that hard to set up a handful of subagents that all use the same initial context and to load that context with what actually matters? The APIs certainly support it.)

Re: GLM-5.3-Flash

#250

Earlier quoted context omitted.

~$4000 USD each on Amazon, $175 for the cable.

The cables are ~USD $50 from AliExpress although I'm not sure what the tariff situation is for Americans (I think I paid $75 all-in CAD for them)

Big "depends". Ali has "ships from China" and "ships from US" stuff. In a lot of cases, though, Amazon also tends to have cheap Chinese knockoffs for a comparable price (although sometimes their product ranking buries them and/or promotes the more expensive knockoffs)

Edit: Yeah I see an ONTi QSFP56 on Amazon for $45, 10Gtek QSFP112 for $62

Post reply on HN