Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

211–220 of 298 posts

Re: GLM-5.3 is now open-weight

#211

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

It IS crazy to drop big money on any AI rig right now imho... the size of models and the cost to run them is falling through the floor as we speak. I'm happy with all of the competition in the APIs on openrouter... I watch that like I used to watch the stock markets, lol. It's great fun.

[deleted]

Re: GLM-5.3 is now open-weight

#212
Interesting model. It's nice to see open-weight models catching up with the top proprietary ones. The coding benchmark results are pretty impressive. Overall, it's great to see progress moving forward. Curious to see how it performs in real-world usage.

Re: GLM-5.3 is now open-weight

#213
post #63

Earlier quoted context omitted.

Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.

I tried using deepseek v4 flash with OpenRouter. It switches between providers too eagerly which resets the cache. Then, each provider begins to rate limit me for providing so many uncached tokens, so it just keeps on switching providers. I'm paying for every token... why rate limit me? It was unusable compared to just using the official Deepseek provider which has a much better cache rate.

You really need to select your provider with Openrouter to get the best experience.

https://openrouter.ai/docs/guides/routing/provider-selection

Re: GLM-5.3 is now open-weight

#214
post #75

Earlier quoted context omitted.

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years. However there are other reasons (e.g. privacy) that might make it worth running locally for some people.

[deleted]

Re: GLM-5.3 is now open-weight

#215
post #154

Earlier quoted context omitted.

I couldn't find any article that states z.ai has a data center in Singapore. There are stories of their new 1 GW data center in China though. Also, openrouter lists HQs on their providers page which matches the regions. https://openrouter.ai/providers

Interesting. I wonder where they're getting their data from then, because they list Z.ai under Singapore, but everything I'm finding says they're based in Beijing. Same with MiniMax.

Many Chinese AI companies have their HQs in Singapore because of US sanctions. If I remember it correctly Manus is also based there. But it's a Chinese company through and through as we know from the events that unfolded after Meta's buyout attempt.

Re: GLM-5.3 is now open-weight

#216

I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack. We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.

What coding harness are you using? I am trying to take the plunge and wondering which I should use

Re: GLM-5.3 is now open-weight

#217

Earlier quoted context omitted.

> I think the biggest reason is to own the stack so your model can't be changed out from under you, The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments. As long as there is demand for a model, it will be hosted by multiple providers.

What if the model is hopelessly obsolete, and thus no demand, but I want that specific model? Owning the weights and hardware is not just solving for one problem. It eliminates all the classes of problems that occur outside of your building, if you have a solar and battery setup. Also, on a more practical basis, what if the way it's served is bad. Maybe I want my specific KV setup, or ultra low quant for entertaining…

> What if the model is hopelessly obsolete, and thus no demand, but I want that specific model?

You can still find a lot of old and completely outdated models on OpenRouter. The providers can scale serving of models up and down as demand arrives, so models don't generally disappear. They're just kept in the mix and the clouds will allocate hardware to it if someone is willing to pay.

In the odd case that it disappears completely, buying the hardware 2 years from now is probably going to be a better deal. That wasn't true if you selectively check the time period before hardware got expensive, but as new hardware comes out we're going to start seeing Strix Halo and old Apple hardware hit the market as people upgrade. It's already happening.

There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities. If you fit that description then there's nothing anyone can say to discourage you from buying your own hardware, but for everyone else I do not recommend buying hardware to self-host LLMs just to save money. I self-host and run a lot of tokens through my setup (non-coding work) but I'm not really saving money.

Re: GLM-5.3 is now open-weight

#220
post #25

I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?

I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.

I agree that they are historically important, but if you want to query old thinking in 2070, you'd probably be better off having a modern model analyze archive.org. If that ever goes down, we're sunk.
Post reply on HN