Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

91–100 of 296 posts

Re: GLM-5.3 is now open-weight

#92

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

Its reasoning leaves a lot to be desired :(

Though I appreciate how good it is at "solid" grunt work and at that price (in fact I am paying the grandfathered subscription price; mostly).

I am planning to let go for my Claude AI subscription which I now use only for "planning" and maybe use that via Open Router as PAYG (at to try how it ends up). But god glm is bad at "talking" and "responding" anything prose. Not only quality but it's almost impossible to tune it and make it let go of its habits and biases and enthusiasms which often result in too many too and fro.

So I sometimes wonder at what point that starts becoming the cost and mental hassle. Maybe it's not there for me yet.

Re: GLM-5.3 is now open-weight

#93

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

So far I don’t regret buying an M1 Max device with 32Gb of RAM. The models available for it keep getting better (running just about okay for interactive use) and 400 GB/s of bandwidth is still considered a lot.

The models are currently improving much faster than the hardware and this doesn’t seem to have plateaued yet.

Re: GLM-5.3 is now open-weight

#95
post #43

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.

I have a dual epyc + 1TB RAM. I could push glm 5.2 to 7 tok/s CPU only.

Re: GLM-5.3 is now open-weight

#96

Earlier quoted context omitted.

[flagged]

Not every tech worker is making top-tier US salaries. For some (I suspect not few) people on HN that $20,000 Mac is almost a year's salary.

and even if you were making such a salary, the quesiton of if the investment on hardware to run llm's locally is still a big if, its OK if you buy the HW cause you'll use it and you get the extra capability as a nice extra, but doesnt make sense to spend so much when you could just get 200$ subs with almost infinite SOTA tokens a month etc (if you dont need the local/privacy aspects of it)

Re: GLM-5.3 is now open-weight

#97

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

That's basically the question I'm trying to answer.

If you're paying Anthropic or OpenAI to use their models, harness, governance, etc., I could see the local inference potentially coming out ahead. They're already starting to ratchet down what your money gets you on their platforms, and that can be expected to continue as the leaders of those companies continue to seek the road to the El Dorado that is being a trillionaire.*

If you're looking to get into the guts of AI development instead of having it handed to you by a provider, that's where it gets murky. I'm wanting to write some sort of agent that does things and get into making outputs consistent in the like, and I'm not sure whether to host something on GCP or buy an M5 Mac.

*Note: El Dorado is a mythical city and many people died trying to find it.

Re: GLM-5.3 is now open-weight

#98
post #29
post #11

h/t to DeepInfra for being the first 3rd party provider for it on OpenRouter ( https://openrouter.ai/z-ai/glm-5.3?endpoint=b711bea7-3994-49... ).

on my TrustedRouter: z-ai/glm-5.3: also Z.ai, Novita, Atlas Cloud, IO.NET

I have seen you advertise your website a few times. I like the idea of not having to trust the router, so I took some time out of my day to critique your website: https://files.catbox.moe/v68cf7.png

My visit to your website went like this:

1. Visit models page

2. Try to find GLM-5.3-Flash (which is among the ~5 models that 90% of people currently care about)

3. Give up scrolling (which would have taken OVER 50 SCROLLS!!!) and use Ctrl + F

4. Try to find input/output/cached price

5. Scroll all the way up to find out which column is what

6. Notice that output price is cut off

7. Notice that the scroll bar is over 100 scrolls further down the page

8. Use Shift + Wheel to scroll horizontally (most visitors probably won't know this trick)

9. Notice that cached price is missing

10. Conclude that this is probably not a serious offering and bounce

There are probably more issues later on, but this is how far I got.

I would suggest you to:

- Deslopify all pages that a user may visit before conversion

- List important models first (see OpenRouter rankings)

- Move the most important information (model name/input/output/cached price) to the left

- Disaggregate the prices per provider (maybe subtables per model? not sure)

- Measure cache hit rate and compute effective price per provider (see OpenRouter)

(- Optional: Fix the broken link on your HN profile page. Currently, the only way to get from this comment to your website is a search engine.)

Re: GLM-5.3 is now open-weight

#99

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

[deleted]

Re: GLM-5.3 is now open-weight

#100
post #25

I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?

I think it would be an important historical document as well. We are potentially looking at the dawn of AGI and one of the most important models ever created. Each model is also a kind of ultimate time capsule, containing a snapshot of the entire human collective mind. If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.

> If you wanted to ask a 2002 person what they thought about future historical events you can just ask them directly.

The weights arent the truth tho, maybe a timecapsule-vhs but i wouldnt trust llm weights more than more hardcore deterministic media that might get preserved to infer facts from an era.

The companies doing the training are becoming the "winners" that are "rewriting history" as they train their models.

Post reply on HN