You guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you…
Chinese laws are not valid in the EU
GLM-5.3-Flash
171–180 of 605 posts
Re: GLM-5.3-Flash
#172so is it confirmed if this is the mysterious OxAlpha model?
Re: GLM-5.3-Flash
#173Earlier quoted context omitted.
> I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc. Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly. There are a lot of…
I have exactly the same opinion Over the last couple years I’ve had to learn sales and understand the thought process behind this better, and I think I’m beginning to understand it The psychology is that most people aren’t really trying to optimize for productivity (even most people who think they are) on an ROI basis, because their compensation is too decoupled from their actual raw output, and more closely coupled…
Re: GLM-5.3-Flash
#174> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini) And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet…
Re: GLM-5.3-Flash
#175Earlier quoted context omitted.
There's soooo much by way of experiments, explorations, tinkering, and even projects that you can't possibly pursue through a some SaaS API. The more reasonable comparison is against rented GPU's, while looking at tradeoffs in latency and upload/download/storage/instance management overhead. Buying hardware for local models is meeting a wholly different need than buying tokens through OpenRouter or whatever.
It cuts both ways. A GPU in your basement is a depreciating asset with fixed computing power and consumes electricity. Switching model providers is trivial.
All decades prior and up to about a year ago, I would have agreed with you. My Framework Desktop, however has appreciated in value by 75% since I bought it. Will it stay there for a long time? Probably not. But it shows that there are no hard and fast rules about things anymore.
Re: GLM-5.3-Flash
#176> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini) And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet…
https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...
Re: GLM-5.3-Flash
#177Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
Re: GLM-5.3-Flash
#178Earlier quoted context omitted.
Well, Luna debuted with 5x higher pricing than is currently available. With the pace of recent development these models might not be relevant by Thanksgiving.
Of course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are…
I don't have enough metrics to compare those costs but still Chinese models have been cheaper except against Luna for me.
FWIW, Luna does everything so well, I just keep using it for all my agents by default.
Re: GLM-5.3-Flash
#179Earlier quoted context omitted.
I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads
The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.
Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.
(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)
Re: GLM-5.3-Flash
#180Earlier quoted context omitted.
Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
Has there been any confirmation about what that model even is? Edit: Ah: > This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.