Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

331–340 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#331
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

> This is absolutely still shy of Sol and Fable, but only just by a hair

What's crazy is that this is a relatively small model - approx. 750B total, 40B active params, while Sol and Fable are one or two tiers above that (Kimi 3 and Qwen 3.8 also ~3T params).

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#332
post #132

Earlier quoted context omitted.

I have already begun winding down my spend on claude and OAI to make room for infra budget. Anecdotal, but I have no doubt a lot of others are doing the same, I very much agree the US players have major issues looming. What an exciting time to be alive!

Not exciting for anyone directly or indirectly invested in a frontier lab or its partners. And that is a lot of people, including you.

Sometimes fear can be quite exciting.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#333
post #70

Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it. Original images: https://image.non.io/neonRamenDesigns.webp GLM 5.3 build: https://html.non.io/neonRamenGLM5.3 Opus 5 build for comparison: https://html.non.io/neonRamen For having no visi…

Blindsight!

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#334
post #289

Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the…

> Anyone using them now? You're gonna have a had time getting straight answer to that out of the internet. There are now 4 different flavours of the Max plan floating around (Legacy V1, Legacy V2, New plans, and the current credit ones). And on top of that they have peak times. So ~8 scenarios, 24 in total across all feedback for their coding plans. So when someone tells you they're having a good time on a GLM coding…

I have V1 Max, and I think they throttled me for using it too much. I was maybe abusing it, by sending out 8 or 16 review agents at a time.

I haven't tried it in a few months, but it went from amazing to unusable really fast.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#335

Earlier quoted context omitted.

does this suggest 5.3 is the same # of parameters as 5.2?

which is the bigger headline that people don't realize. this is 744b and its head to head with Kimi K3 (2.8T), smashes DS v4 pro (1.5T). even Opus and Sol are rumored to be 1.5T+ this is half the size!

in terms of performance the formula seems to be : dense parameters = sqrt(total*active)

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#336
post #310

Earlier quoted context omitted.

They're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast. If it won't attack my stuff, it won't help me build my stuff to be secure.

Exactly my CoT! I hope z.ai won’t change this behavior after training it on our input the same way as Anthropic did (shame on you, folks, seriously)

> after training it on our input the same way as Anthropic did (shame on you, folks, seriously)

What do you mean with this? Honest question!

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#337

Earlier quoted context omitted.

> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.

Impossible with source code, possible to bypass with app/site

Don't we already do this with services like Let's Encrypt, which is arguably more sensitive? If you had the codebase you could fake it, but it would still provide some amount of protection against abuse.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#338

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

You should try a better harness. Try pi, or ohmypi if you want a good OOB experience

I’m in the Claude code harness for everything boat too. What are the alternatives?

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#339

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

You should try a better harness. Try pi, or ohmypi if you want a good OOB experience

what is this comment based on ? vibes?
Post reply on HN