Live data from Hacker News

GLM-5.3 is now open-weight

huggingface.co

101–110 of 296 posts

Re: GLM-5.3 is now open-weight

#101
post #25

I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?

There’s not such a straightforward relationship between safety and model sis. According to the book The Thinking Game, lower quality models at that time were considered less safe, because they could be easily tricked into doing harmful stuff. In the book, Dario (of Anthropic) was the head of safety at openAI and was responsible for pushing for 10x scaling in training to make the models safer . It does make sense, a s…

> According to the book The Thinking Game...

According to me, this is nonsense.

Re: GLM-5.3 is now open-weight

#102
post #63
post #11

h/t to DeepInfra for being the first 3rd party provider for it on OpenRouter ( https://openrouter.ai/z-ai/glm-5.3?endpoint=b711bea7-3994-49... ).

Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.

I just checked because I was a bit paranoid, but I have a 96.6% cache hit rate for GPT-5.6 Luna and 96.8% for Opus 5.

Re: GLM-5.3 is now open-weight

#103

GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better. Assuming you’re willing to drop a fat…

When we consider: * LLM usage is new for the world * Models are evolving quickly with high worldwide competition * Hardware is evolving despite RAM shortages Is investing a huge sum of money in equipment for local inference a wise use of money? Or are M5 Ultra and equivalently priced local inference hardware future-proof enough to be worth it relative to how the market is evolving? Maybe it’s all a question of what y…

Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally. It’s the same reason my entire development stack from editor to compiler is open source. I don’t need to modify it today, but I always must retain the option to do so later.

There are several things I do in my life that only pay off in the event of a big disaster, like an extended internet outage, civil unrest, supply chain disruption, war, etc.

I like to be able to do the things I do even if offline for weeks.

I spent a lot of money for more flash in my iPad Pro so I can keep all of offline wikipedia and OSM in it, for example, along with tons of books. It’s sort of like being a digital prepper. (Being a prepper is a spectrum, from anyone who keeps food in their pantry to people building bunkers under their house - how much you invest is a personal prudence and threat modeling decision.)

Also, privacy. And when I got the Mac Studio the 512GB was only $15k, which is dirt cheap for that much VRAM.

Re: GLM-5.3 is now open-weight

#104
post #63
post #11

h/t to DeepInfra for being the first 3rd party provider for it on OpenRouter ( https://openrouter.ai/z-ai/glm-5.3?endpoint=b711bea7-3994-49... ).

Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.

I tried using deepseek v4 flash with OpenRouter. It switches between providers too eagerly which resets the cache. Then, each provider begins to rate limit me for providing so many uncached tokens, so it just keeps on switching providers. I'm paying for every token... why rate limit me? It was unusable compared to just using the official Deepseek provider which has a much better cache rate.

Re: GLM-5.3 is now open-weight

#105
post #2

I've been using it more and more. Feels like Opus 4.8, in the best possible way.

I'm starting to think Opus 4.8 is significantly smaller than most people assume. If it's significantly larger than GLM 5.3 (I've heard some insane guesstimates out there like upwards of 5T params or more), that would prove rather embarrassing for Anthropic.

I think the Western labs are burning through funding and compute to maintain the lead at any cost, efficiency be damned.

Re: GLM-5.3 is now open-weight

#106
post #36

Earlier quoted context omitted.

They already publish gpt-oss which is several generations better than gpt-3

GPT-3 is a different model than gpt-oss and is therefore not an answer to the question. I cannot stand using gpt-oss, but I miss some of the creative spark of GPT-3 davinci dearly.

I’ve found a very similar creative spark with Gemma4 base models. You’ll have to do old school prompting, but it’s kinda fun too.

Re: GLM-5.3 is now open-weight

#107
post #90

Earlier quoted context omitted.

Sort of depends on how well the core reasoning works. It’s not a big effort to connect an LLM to a search provider. You do pay for the tokens, but in theory on a smaller model each token is cheaper.

honestly using search isn't that great, you mostly get SEO slop, it usually won't help the model ask the right questions

Try Parallel.ai (no affiliation). Instead of keywords, the model writes objectives.

Re: GLM-5.3 is now open-weight

#108

Earlier quoted context omitted.

There’s not such a straightforward relationship between safety and model sis. According to the book The Thinking Game, lower quality models at that time were considered less safe, because they could be easily tricked into doing harmful stuff. In the book, Dario (of Anthropic) was the head of safety at openAI and was responsible for pushing for 10x scaling in training to make the models safer . It does make sense, a s…

> According to the book The Thinking Game... According to me, this is nonsense.

I mean, none of this is controversial from a historical perspective, people did think that way. Whether they were right about us is another matter.

Re: GLM-5.3 is now open-weight

#110
post #102
post #63

Earlier quoted context omitted.

Have you guys been having a good experience with OpenRouter? I tried it out recently with Claude, and it cached no tokens, charging me $200 for one conversation of 11 messages.

I just checked because I was a bit paranoid, but I have a 96.6% cache hit rate for GPT-5.6 Luna and 96.8% for Opus 5.

Hm, thanks, it must have been some OpenWebUI bug, thank you.
Post reply on HN