Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

321–330 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#321
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

Wordpress having a high number of vulnerabilities not surprising lol

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#322
post #70

Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it. Original images: https://image.non.io/neonRamenDesigns.webp GLM 5.3 build: https://html.non.io/neonRamenGLM5.3 Opus 5 build for comparison: https://html.non.io/neonRamen For having no visi…

wow, what kind of stuff does that script do? I've seen non-vision models analyze images, but mostly histograms, color averages etc. This one seems to actually understand the image itself and reproduce the layout, very impressive

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#323

Earlier quoted context omitted.

> It's what enron was doing; Enron hid billions of dollars in debt and fake profits. Is this what you think is happening here?

Nvidia has made a lot of very suspicious circular funding deals. I suspect we’ll find fraud when the bubble bursts yes.

You're the first person I hear claiming NVIDIA is hiding billions of dollars in debt and fake profits, never mind at the scale of Enron.

Bold claim!

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#324

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

You should try a better harness. Try pi, or ohmypi if you want a good OOB experience

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#325

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.

Impossible with source code, possible to bypass with app/site

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#326
post #5

This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…

> this is just GLM 5.2 with post-training magic Isn't post-training turning out to be the most important part?

It basically has been ever since they started using RLVR for reasoning (esp. coding & math), with the DeepSeek-R1 paper being what let the cat out of the bag.

The Gemini 3.7 Flash model released yesterday, and all the 3.x Flash models, are still based on the Gemini 3 pre-training run from January 2025 !!

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#327
post #161

Earlier quoted context omitted.

This is going to be catastrophic. Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.

I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain ma…

Thanks for sharing.

Is lack of inference chips due to the trading blocks by trump administration? What if Trump agrees to sell chips to china, would they collapse then? That's not a very strong position to be at

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#328

" a judge agent then attempts each task to verify that it is actually solvable " I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.

[deleted]

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#329
post #258

Earlier quoted context omitted.

This is going to be catastrophic. Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.

Why is anything going to be catastrophic? Companies can go bankrupt without catastrophes for the rest of us. Happens all the time.

I read it as catastrophic for the companies trying to IPO. It'll be great for the rest of us though.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#330

Earlier quoted context omitted.

That makes them at best temporary middlemen. It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't). Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.

The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.

If user data would become such a key ingredient (which it might, i actually remember noam shazeer talking about the importance of user data), i think chinese labs can still get it from china, as keep in mind it ahs a billion people behind the great firewall banned from using us llms. And btw broadly for any gap like this, you really gotta consider that if its becoming a bottleneck, chinese labs will find a way to buy it from one of the labs unless theres strict regulation at the government level
Post reply on HN