Live data from Hacker News

GLM-5.3: Frontier coding with emergent cyber capabilities

z.ai

461–470 of 626 posts

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#462
post #128

Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/ Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high. I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Pro…

> I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago? You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities. Personal…

I find that LLMs also generate a lot of false positives, or extremely minor issues that don't warrant a fix (that are always overstated by the LLM as very important!). Signal to noise is still not great and requires somebody to wade through and pick out the actual good findings.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#463
post #197

Earlier quoted context omitted.

They are selling it, not running it. It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers. I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component.

I take it we saw different demos. I'm under no NDA, if you actually want to know what's up.

I'm assuming you're alluding to them selling a managed solution, alongside the unmanaged solution that the GP is referring to?

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#464

Earlier quoted context omitted.

thank you all! got something to tinker with this weekend i like to challenge my assumptions and try new tools

Just as a +1 anecdote. I enjoy using pi a lot. I used to h think the harness matters a lot but with the current iteration of models I am starting to sway that while it matters it’s less and less important and that CC is bloated. I did some quick tests when I switched and a task that would take $5 in tokens would be completed in $0.50 in pi. Very anecdotal and I don’t have a test framework setup to make this very offi…

are ohmypi and pi related?

that's a very compelling use case, thank you

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#465

Earlier quoted context omitted.

Can OpenCode dispatch background subagents yet? I tried it a week ago and saw nothing. This is 99% of my workflow at this point.

In the new v2 beta, yes. Major QoL upgrade, so much less sitting around waiting.

The v2 branch of OpenCode has not been touched for months, if it's beeing developed then I don't know where.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#466

I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…

> I understand that such models can be used by malicious actors, but it’s fair to have it publicly available I feel like there should be some mechanism to prove you own the code/app/site/whatever and it will remove the guardrails from the LLMs allowing them to find and fix these vulnerabilities.

This is a "they have guns so we need guns" scenario.

You can't guarantee everyone else will use a neutered model.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#467

Earlier quoted context omitted.

What the person above is suggesting: * https://pi.dev/ * https://omp.sh/ (no personal opinions of either, links might be useful) I think that OpenCode is nice, their CLI version is enjoyable and their desktop/web version is okay : * https://opencode.ai/ I also quite like driving OpenCode through something like Kepler / Paseo and tools like that (with those I can still use my Anthropic Condition by Claude Code being t…

Last time I tried some of these, none of them had the "manual mode" that CC has, where it shows you change by change as diffs and you can edit them before accepting and moving on to the next change. I like that because if it's going off pattern I can spot it early on and guide it correctly, instead of having to review the whole completed diff at the end when it's too late. I should spend the weekend checking them out…

The philosophy with Pi is it is minimal (but functional) out of the box and easily extensible. I'm not familiar with that feature but I would not at all be surprised if someone already coded a Pi extension that does it.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#468

Earlier quoted context omitted.

> That was my thought too. For all of Anthropic's talk about their "adversaries" It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.

This is a coherent explanation for why federal model censorship has started with cyber capabilities. But this GLM model release is an in-your-face challenge to that policy. They now have to either set models free or impose a censorship regime that will put anyone not under it at an advantage. Or muddle along in the middle as usual.

> Or muddle along in the middle as usual.

I'm not a gambling person, but if I was this would be my bet.

Re: GLM-5.3: Frontier coding with emergent cyber capabilities

#469

Earlier quoted context omitted.

Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.

> Anyone who knows anything realises banning things is a) impossible and Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which make…

This is different now. US labs and companies are not releasing frontier-level models openly (specially those capable of assisting cyber intelligence work), but commercializing them instead. Thus, any ban would not be symmetrical to begin with, and that is precisely what maintains the balance.
Post reply on HN