Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

291–300 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#291

Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3? I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model... Nowadays, I’d probably go with their Max plan if the…

I found with GLM I was better off using plans from either Neuralwatt or Ollama.

But Neuralwatt significantly raised their rates since then.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#292
" a judge agent then attempts each task to verify that it is actually solvable "

I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#293
post #125
post #105

Earlier quoted context omitted.

I can’t come up with a use case where I couldn’t extract the image details using another, multimodal model and pass it into the GLM’s context with as many details as I need.

I think you lose a lot by not having the vision capability shared with the text. It is the joint reasoning across them where the power lies (the same model that sees the code and made the changes to produce the visual presentation, sees the image of it and reasons about it).

I mean this is assuming the thing you're working on has a UI? Not all of us work in that space.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#294

Earlier quoted context omitted.

The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so? [1] https://www.washingtonpost.com/graphics/2020/world/national-...

You should probably realise that the world has radically changed since then. This kind of thing works when you have a significant lead in the field that makes keeping vulnerabilities open sufficiently low risk for your own side. But if your adversaries have similar capabilities, then the calculation changes.

Has anything changed? Governments are hoarding undisclosed vulnerabilities, using them as they see fit instead of fixing. Every espionage, surveillance, or war campaign (see Russia v Ukraine, US/Israel v Iran etc) is followed by a ton of burned 0-days.

>This kind of thing works when you have a significant lead in the field

No? It works even if the adversary has the same capabilities. It only stops working when everything is fixed.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#295

Earlier quoted context omitted.

OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize. These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin. I just don't see how you justify a trillion valuation for US AI labs when the underlying models are…

Another interesting potential market here will be 'LLM in a box'. All the hardware and other tooling in a prebuilt, but modular, package ready to go. Pay one up-front cost, get a system running [whatever open LLM] with a token rate of [x], optionally configured to be immediately ready for distributed usage. Basically the opposite of cloud stuff: no rent, no dependency, 100% guaranteed uptime, guaranteed security/priv…

I see ads for this all the time.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#297

Earlier quoted context omitted.

That makes them at best temporary middlemen. It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't). Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.

The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.

But is that data good? That's the question. As in, is my usage at work:

a) indicative of problems that aren't already out there in the wild? (no) b) are the responses I'm getting so good and novel that the model can improve itself? (no)

It's the garbage in garbage out idea, just scaled up. If the model gave a bad answer, and I didn't catch it, and you now train on that I/O pair (my perhaps crappy prompt, the bad output), then you're not going to improve anything.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#298
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

> ... Anthropic's Project Glasswing is supposed to find them quite a while ago? That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.

> and that Apple would not attribute to GLM

That was a wtf to me, so I checked Apple’s latest iOS release security content and GLM & z.ai is mentioned once (under WebKit), Anthropic is mentioned twice, Codex is mentioned once. Not clear if there are other instances where the model did most of the work but wasn’t credited. I didn’t bother to check other releases.

https://support.apple.com/en-us/128066

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#299

Earlier quoted context omitted.

The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.

I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this. Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volum…

I agree. And even if they were able to do it for one more round, it's not a sustainable strategy. What they (Anthropic and OpenAI) need to do is build platforms and integrate verticals.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#300
post #181

Earlier quoted context omitted.

There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.

Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.

Unfortunately, those in power, pretty much all over the world, lie/deceive themselves and believe they can.
Post reply on HN