This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
> This is absolutely still shy of Sol and Fable Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they…
GLM-5.3: Frontier coding with emergent cyber capabilities
491–500 of 626 posts
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#492Earlier quoted context omitted.
> ... Anthropic's Project Glasswing is supposed to find them quite a while ago? That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
> That was my thought too. For all of Anthropic's talk about their "adversaries" It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#493Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#494Earlier quoted context omitted.
Based on the fact that Claude Code is only optimized for Anthropic models, whereas Pi and Omp are optimized for a wide variety of models, including open weights.
they are not really optimized for 'wide variety of models' . what optimization did pi do for glm 5.3?
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#495Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#496I bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#497Earlier quoted context omitted.
Why should I apply for *cybersecurity* approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, "programming") somehow is cybersecurity now?
Your tools refusing to do your bidding is an absurd idea in the first place Imagine asking for permission to use your hammer
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#498Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#499This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice. How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with p…
Re: GLM-5.3: Frontier coding with emergent cyber capabilities
#500Earlier quoted context omitted.
You should try a better harness. Try pi, or ohmypi if you want a good OOB experience
what is a harness? The comments below are mixing IDE/ADE but other suggestions are purely terminal things and I don't get what their value is over just a terminal. Is a harness like a loop where it's just a vague thing that everyone nods about but everyone is nodding at something different?
The tool calls will be, among other things, something like ReadFile, RipGrep, PatchFile, Shell.
When people talk about the value of different harnesses, they're also implicitly talking about the quality of the system prompt.
The same exact model, when given a different set of tools and a different system prompt, can behave differently.