Earlier quoted context omitted.
It's not that GLM5.3 in full precision unquantized is any smaller, it's 141 * 5.4GB files at approx 770GB which is about the same size as 5.2.
Hold on.. the routed experts are in FP8 now? Previously they were in BF16. Nice, this shall cut my download time by half!
GLM-5.3 is now open-weight
121–130 of 297 posts
Re: GLM-5.3 is now open-weight
#122I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack. We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
Re: GLM-5.3 is now open-weight
#123Re: GLM-5.3 is now open-weight
#124h/t to DeepInfra for being the first 3rd party provider for it on OpenRouter ( https://openrouter.ai/z-ai/glm-5.3?endpoint=b711bea7-3994-49... ).
on my TrustedRouter: z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET
Re: GLM-5.3 is now open-weight
#125I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?
There’s not such a straightforward relationship between safety and model sis. According to the book The Thinking Game, lower quality models at that time were considered less safe, because they could be easily tricked into doing harmful stuff. In the book, Dario (of Anthropic) was the head of safety at openAI and was responsible for pushing for 10x scaling in training to make the models safer . It does make sense, a s…
In the times of GPT-3 I'd scoff at the idea of an LLM doing any hacking; today, I'm running several AIs on my code before publishing, and they are finding (and demonstrating!) RCEs on my localhost server.
For example, one found a missing check in a third party JWT library which allowed full account takeover, which I'd have never even looked at.
Hence I don't believe a single word coming out of these people's mouths. Their "beliefs" are just marketing.
Re: GLM-5.3 is now open-weight
#126I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack. We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
I assume that's about 5.3 Flash, not full?
Re: GLM-5.3 is now open-weight
#127Does this mean it'll be on Bedrock soon? I hear great things about this model but I want AWS data handling practices...
Re: GLM-5.3 is now open-weight
#128A stake through Amodei's heart.
Re: GLM-5.3 is now open-weight
#129Earlier quoted context omitted.
When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…
Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. The object permanence of not having to reinvent the world every time a model gets sunsetted has value.
with open models, there is ecosystem/market of providers, where you can easily switch to provider you like
Re: GLM-5.3 is now open-weight
#130Earlier quoted context omitted.
When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…
Tools vs services in my mind. There is no guarantee any provider will continue to do what they are doing for you at the price they are doing it. The object permanence of not having to reinvent the world every time a model gets sunsetted has value.