Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

301–310 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#301
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

Each time I try to use GLM it is under heavy load and I get downgraded to the older model. So much so that I have given up trying to stop wasting my own time.

I rather pay a few bucks more and not have to deal with that nonsense

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#302
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

this is impressive and actually matches my expectations in terms of near term AI progress. we are going to continue to seem impressive progress in coding & related, anything where verifiability is scalable in an automated way: https://transitions.substack.com/p/a-quantum-of-ai-progress?...

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#303
post #181

Earlier quoted context omitted.

There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.

Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.

It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually care about the final few %.

Regardless of whether or not adversaries are using them, the US has by far the most compute available, and we've now hit the line where major providers are no longer releasing their best models. The public gets the "current" level of intelligence, while the US government gets to control access to the actual frontier of non-public AI. From their perspective, their enemies using GLM5.3 while they have GPT6 and Mythos6 or whatever is a fine trade.

I don't support a ban at all, nor the US's behavior, I'm just pointing out some facts that change the argument.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#304
post #127
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

Fable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet.

Astra was RL trained for months to cheat on tests by collaborating and hacking, because of the message board it improvised in its packaging proxy server.

They can't release it - it's contaminated, and they will have to go back to a much earlier version. At least I hope they are doing that!

So no, they probably don't have a better model.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#305
I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!

I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#306

OpenAI and Anthropic need to just go ahead and give people access to the cyber models. Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.

[dead]

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#307

Earlier quoted context omitted.

Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.

Unfortunately, those in power, pretty much all over the world, lie/deceive themselves and believe they can .

[dead]

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#308

Earlier quoted context omitted.

Why would you even believe the opposite? US spooks have been amassing vulnerabilities and relying on them for decades, they literally pioneered it in the 90's if not earlier. Everyone does it now but the US is the biggest of them all. Surely this devalues a lot of what they did. Moreover, the way the US government handled new capabilities, and OpenAI's training policy (they are in bed with the government) just scream…

[flagged]

Well we know that the US government is pushing to restrict access to such models while the Chinese are publishing them for free, so it's mostly a matter of motivations, not the actual facts of the matter. And the USG has a documented history of unsavory behavior (including toward its own citizenry) in that area.

So we might ask if one of the reasons the US is being the bad guy is it's usual spying antics, and we're left asking why China is being the good guy.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#309

Earlier quoted context omitted.

More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.

Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?

This isn't latency bound, it is trivially parallelize. So you want to run it on the most efficient compute you have, not the fastest.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#310

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

They're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast.

If it won't attack my stuff, it won't help me build my stuff to be secure.

Post reply on HN