Earlier quoted context omitted.
> get myself four sparks at a decent price Wow, if you don't mind me asking. How and where?
I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
GLM-5.3-Flash
241–250 of 605 posts
Re: GLM-5.3-Flash
#242Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is. What irks me about this is that the harnesses seem to be just an afterthought here. Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, e…
Re: GLM-5.3-Flash
#243Earlier quoted context omitted.
I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plan…
This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does. Not…
Or just rent something substantial for like $4/hr on runpod or w/e to do that.
My gripe is this persons compute is wasteful and makes it harder for me to buy something with like 64gb ram to do normal work and run containers while I keep using cloud models.
Someone else calculated the break even being 10 years, it’s just dumb. And I think it’s clear there won’t be a big rug pull anymore, there are too many open models and providers now.
Re: GLM-5.3-Flash
#244Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experiment…
Re: GLM-5.3-Flash
#245Earlier quoted context omitted.
> DeepSeek token prices are continuing to _increase_ One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...
You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
Re: GLM-5.3-Flash
#246On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M How is the business model of Anthropic/OpenAI will sustain?
Re: GLM-5.3-Flash
#247Earlier quoted context omitted.
I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads
Checked a couple days ago and looks like we're at about 3.5x 2020 memory prices (looking at just $/GB).
Re: GLM-5.3-Flash
#248Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost…
I don't know how anyone can actually use Luna max on ANY real workload. I've had Sol orchestrate a bunch of Luna agents, these agents were explicitly given small chunks of larger objectives and they still filled their entire context windows with just reasoning tokens, until compaction hit, and then reasoning again. I've probably wasted a good 40% of my weekly usage on Luna Max agents just thinking and not writing a s…
I’m pretty sure that plain Sol, serially, could have finished the task faster, cheaper, and far more accurately. I’m also pretty sure that any competent subagent orchestration could have gotten it done with even very simple subagents quickly and cheaply.
(Is it really that hard to set up a handful of subagents that all use the same initial context and to load that context with what actually matters? The APIs certainly support it.)
Re: GLM-5.3-Flash
#249Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.
Re: GLM-5.3-Flash
#250Earlier quoted context omitted.
~$4000 USD each on Amazon, $175 for the cable.
The cables are ~USD $50 from AliExpress although I'm not sure what the tariff situation is for Americans (I think I paid $75 all-in CAD for them)
Edit: Yeah I see an ONTi QSFP56 on Amazon for $45, 10Gtek QSFP112 for $62