Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…
> ... Anthropic's Project Glasswing is supposed to find them quite a while ago? That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
GLM-5.3: Frontier coding with emergent cyber capabilities
311–320 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#312Earlier quoted context omitted.
Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually care about the final few %. Regardless of whether or not adversaries are using them, the US has by far the most compute available, and we've now hit the line where major providers are no longe…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#313Earlier quoted context omitted.
Meanwhile the GB300 used by hosted llms: GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional https://pi3g.com/nvidia-gb300-specifications-including-memor... If you think M7 will hit even 15% of these speeds you're very optimistic.
A hosted instance serves multiple customers at a time. A local model only one.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#314This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
Have you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer step-by-step—hiding the memorization from the user. 0: https://www.alphaxiv.org/abs/2608.09867?hl=en-GB
for some reason I couldn't find any way to download it from that website.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#315Earlier quoted context omitted.
I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain ma…
That makes them at best temporary middlemen. It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't). Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#316Earlier quoted context omitted.
Circular investment deals and investment deals at valuations which have no possible justification.
Let me ask it differently. You state the companies and their investors are colluding. Who are they colluding against ?
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#317I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…
They're invaluable for developers to fix their code. This is definitely an area where AI decisively beats human devs in a very valuable way. It can try so much surface area so fast. If it won't attack my stuff, it won't help me build my stuff to be secure.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#318Earlier quoted context omitted.
its interesting, as it's typically the banks and against the public at large because the goal is to jimmy up valuations to justify IPOs then sell on opening; just like spacex. It's what enron was doing; it's what most of crypto's offshoots were doing. Sure you can blame the marks of the grift and say "well the public should know they're faking all this cash flow expectation". It seems like you're either driving the g…
> It's what enron was doing; Enron hid billions of dollars in debt and fake profits. Is this what you think is happening here?
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#319I'm not that up to date with the latest AI developments, but I noticed that this article seems to use "Cyber Capabilities" as a shorthand for the model's ability at cybersecurity tasks? Is that now an established expression, same as "crypto" now refers to cryptocurrencies rather that cryptography? Because "cybernetics" actually means something different (yeah, old man yelling at clouds, I know)...
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#320I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…
I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.