GLM 5.2 beats Claude in our benchmarks
semgrep.dev
GLM 5.2 beats Claude in our benchmarks
1–10 of 559 posts
Re: GLM 5.2 beats Claude in our benchmarks
#2After installing, do a `n8 build` to build the image, then `n8 --danger --provider opencode interactive` to launch it in a container.
Signup for GLM-5.2 here: https://z.ai
Re: GLM 5.2 beats Claude in our benchmarks
#3Not that it would make any sense.
Re: GLM 5.2 beats Claude in our benchmarks
#4GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.
US imposing export restrictions on a model from China?
Re: GLM 5.2 beats Claude in our benchmarks
#5Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?
Re: GLM 5.2 beats Claude in our benchmarks
#6Secondly these are "just" IDORs, arguably the easiest class of vulnerabilities.
Thirdly it compares to GPT 5.5 and Opus 4.8.
No, we don't have Mythos at home.
Re: GLM 5.2 beats Claude in our benchmarks
#7Here, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc. Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?
Re: GLM 5.2 beats Claude in our benchmarks
#8GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.
>GLM export controls incoming? US imposing export restrictions on a model from China?
Re: GLM 5.2 beats Claude in our benchmarks
#9It reads like an ad. Secondly these are "just" IDORs, arguably the easiest class of vulnerabilities. Thirdly it compares to GPT 5.5 and Opus 4.8. No, we don't have Mythos at home.
mythos is 1000% to run inference on a <10% better model would've been very damning.
Re: GLM 5.2 beats Claude in our benchmarks
#10Here, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc. Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?
I think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"