Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

281–290 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#281
post #185

Earlier quoted context omitted.

Do you actually believe this?

The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so? [1] https://www.washingtonpost.com/graphics/2020/world/national-...

You should probably realise that the world has radically changed since then. This kind of thing works when you have a significant lead in the field that makes keeping vulnerabilities open sufficiently low risk for your own side. But if your adversaries have similar capabilities, then the calculation changes.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#282

Earlier quoted context omitted.

The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.

I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this. Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volum…

assuming the technology of model architectures does not gain any further breakthroughs that returns us back to the gains previously seen. I'm of the opinion that we still have some discoveries on the mathematical side of the fence to go that will improve models further.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#283
post #127
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

Fable finished training 6+ months ago. At this point, Anthropic only needs to release models to the public when the competition forces them to. OpenAI also has a better model (Astra) that they haven't released yet.

are you just assuming capitalism will keep burning money to keep ahead?

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#284
post #181

Earlier quoted context omitted.

There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.

Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.

> Anyone who knows anything realises banning things is a) impossible and

Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which makes "banning" maybe not completely useless for certain things.

"Banning" things that can be easily copied over cyberweb transportation pipes feels like an fool's errand though, regardless of what it is. It's just too easy to get around, compared to actual physical items I suppose.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#285
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

> I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?

You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities.

Personally, I've had Fable 5, GPT 5.6 Sol, and GLM 5.2 all looking for correctness issues in an old abandoned WIP codebase of mine and all of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#287

One htought I had; if The chinese allow unfettered access to cyber capabilties while th US does it's best to neuter it's model releases, from China's point of view they have the US all tied up in knots dealing with problems they don't give people the tools to solve. China giggles as it watches the US under threat from people using it's models. The US is restricting citizens from owning this particular kind of weapon,…

I suspect Anthropic wanted the US gov to ban Mythos for marketing. If it turns out to be bad for them, the US gov will likely suddenly unban models.

I suspect mythos demonstrated a fully autonomous offensive hack in a similar way Open AI's models performed, and the government is reacting to it the same way we reacted to blackhat.

The threat is real.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#288

Earlier quoted context omitted.

The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so? [1] https://www.washingtonpost.com/graphics/2020/world/national-...

I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.) I'm curious: to any professional vulnerability researchers reading this, what do you think?

I used to call everything a conspiracy theory, but then Glenn Greenwald published "No Place to Hide: Edward Snowden, the NSA and the Surveillance State".

Now i know that reality is worse than the worst conspiracy theorist.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#289

Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the…

> Anyone using them now?

You're gonna have a had time getting straight answer to that out of the internet. There are now 4 different flavours of the Max plan floating around (Legacy V1, Legacy V2, New plans, and the current credit ones). And on top of that they have peak times. So ~8 scenarios, 24 in total across all feedback for their coding plans.

So when someone tells you they're having a good time on a GLM coding plan it's damn near unusable as a datapoint unless both parties are very clear about what precisely is being discussed

[It's been good for me though...V1 Max off peak...which is basically the best of the 24]

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#290

This will be roughly on pair with Kimi K3, but using a third of its parameters. Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third. Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.

Congrats def in order but as usual the proof will be in the pudding of actually running the thing.

GLM 5.2 has token efficiency problems. It's not a stupid model, but it takes a lot of "thinking" to produce not-stupid results. ("But wait...").

Which makes its pricing deceptive.

I tried to get by through the month of June on just GLM 5.2 and it was ... fine-ish for about two weeks. But the provider situation wasn't ideal.

Post reply on HN