Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

1–10 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#2
You can launch GLM-5.2 in Opencode using Nemesis8: https://github.com/DeepBlueDynamics/nemesis8#nemesis-8

After installing, do a `n8 build` to build the image, then `n8 --danger --provider opencode interactive` to launch it in a container.

Signup for GLM-5.2 here: https://z.ai

Re: GLM 5.2 beats Claude in our benchmarks

#4

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

>GLM export controls incoming?

US imposing export restrictions on a model from China?

Re: GLM 5.2 beats Claude in our benchmarks

#5
Here, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc.

Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?

Re: GLM 5.2 beats Claude in our benchmarks

#7
post #5

Here, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc. Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?

I think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"

Re: GLM 5.2 beats Claude in our benchmarks

#8
post #4

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

>GLM export controls incoming? US imposing export restrictions on a model from China?

While unlikely , it is not without precedent , there are restrictions on ASML a Dutch company to sell EUV machines

Re: GLM 5.2 beats Claude in our benchmarks

#9
post #6

It reads like an ad. Secondly these are "just" IDORs, arguably the easiest class of vulnerabilities. Thirdly it compares to GPT 5.5 and Opus 4.8. No, we don't have Mythos at home.

>Thirdly it compares to GPT 5.5

mythos is 1000% to run inference on a <10% better model would've been very damning.

Re: GLM 5.2 beats Claude in our benchmarks

#10
post #5

Here, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc. Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?

I think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"

Are you replying to a response to the original comment? I looked but i didn't see anyone saying he's throwing shade.
Post reply on HN