Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…
GLM-5.3: Frontier coding with emergent cyber capabilities
321–330 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#322Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it. Original images: https://image.non.io/neonRamenDesigns.webp GLM 5.3 build: https://html.non.io/neonRamenGLM5.3 Opus 5 build for comparison: https://html.non.io/neonRamen For having no visi…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#323Earlier quoted context omitted.
> It's what enron was doing; Enron hid billions of dollars in debt and fake profits. Is this what you think is happening here?
Nvidia has made a lot of very suspicious circular funding deals. I suspect we’ll find fraud when the bubble bursts yes.
Bold claim!
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#324I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#325I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…
> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#326This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
> this is just GLM 5.2 with post-training magic Isn't post-training turning out to be the most important part?
The Gemini 3.7 Flash model released yesterday, and all the 3.x Flash models, are still based on the Gemini 3 pre-training run from January 2025 !!
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#327Earlier quoted context omitted.
This is going to be catastrophic. Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain ma…
Is lack of inference chips due to the trading blocks by trump administration? What if Trump agrees to sell chips to china, would they collapse then? That's not a very strong position to be at
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#328" a judge agent then attempts each task to verify that it is actually solvable " I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#329Earlier quoted context omitted.
This is going to be catastrophic. Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
Why is anything going to be catastrophic? Companies can go bankrupt without catastrophes for the rest of us. Happens all the time.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#330Earlier quoted context omitted.
That makes them at best temporary middlemen. It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't). Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.