Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

41–50 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#41

Earlier quoted context omitted.

> making software better instead of banning it That would be the rational thing to do. > financials and token prices I do not think the government thinks this deeply. Market manipulation might be a rational, if unethical reason to ban open source models. But this admin banned Anthropic models to "own the libs." They will continue to ban what they want for whatever reason they want. I don't think those reasons will be…

Yeah, the current admin is reactionary, they appear to put little thought in, or at least disregard input they dislike. I don't think Ant's ban was about "owning the libs" as much as it was asserting dominance over someone who spoke up counter to the admin's aims and claims. They do listen to money, which is where I see Big Ai paying for executive orders (because the admin forgot what it means to compromise as part o…

[deleted]

Re: GLM 5.2 beats Claude in our benchmarks

#42

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

I think state-of-the-art AI is going to be defense industry only from now on. We can have our toy drones but not the Predators and Reapers.

Re: GLM 5.2 beats Claude in our benchmarks

#43
post #8
post #4

Earlier quoted context omitted.

>GLM export controls incoming? US imposing export restrictions on a model from China?

While unlikely , it is not without precedent , there are restrictions on ASML a Dutch company to sell EUV machines

That’s because the Department of Energy originally funded and contributed IP to the EUV Corp joint venture between several semiconductor companies (including ASML and Intel). Their ability to export control EUV was part of that original agreement that the entire technology is built on.

Re: GLM 5.2 beats Claude in our benchmarks

#44

> [...] beating Claude Code (32%) at roughly $0.17 per vulnerability found Claude Code is an agent harness, not an LLM. Claude is a brand (or group of LLMs), not an LLM.

Yes, and the article author is fully aware of that. Thank you for pointing out this small mistake though.

Re: GLM 5.2 beats Claude in our benchmarks

#45

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

The Americans may ban the use of the Chinese models in America. But like the Chinese car ban, everyone else will use them.

Re: GLM 5.2 beats Claude in our benchmarks

#46
post #37

Apparently GLM 5.2 is 753B parameters [1], what kind of hardware are people using to run this locally? [1] https://huggingface.co/zai-org/GLM-5.2

follow antirez - https://x.com/antirez/status/2071173841175363905?s=20

Thats quantized

Re: GLM 5.2 beats Claude in our benchmarks

#47

Earlier quoted context omitted.

> is it going to be a crime to have them, run them, make them available? Now you're getting it! Commerce will call it a munition and those harboring it as harboring illegal/foreign munitions. No business will take the hit, so they will quickly deplatform the models. No end user has the GPU capacity to use GLM 5.2 or similar models at full precision so the government will call the problem "mostly solved." But they mig…

Or we use the models to work on fixing vulns and stop over-blowing the doom scenarios. Gotta save the kids and kill the terrorists though! I'm for making software better instead of banning it based on what the rich and powerful claim. I suspect the real fear is that open weight models undermine the financials and token prices they thought were going to pay off their ludicrous spending because they have all raced and…

> making software better instead of banning it

We're still in the middle of the cambrian explosion.

If Anthropic was capable of developing Opus 4.49-4.5 2H 2025.... then any company with a research team capable of reading all the papers and press releases will be capable of producing Opus 4.8 by the end of 2027, either raw model competency, or in a harness like claude code (or better with both). I guess what I am trying to say is that Opus 4.5 does not represent the edge of agentic capability, merely somewhere in the thick meaty layer of "functional and achievable".

We can draw the line at Sonnet 4.6 in the US but much like encryption export restrictions in the 1980s, the line drawn will be laughably low within a few years and simply unthinkable in a decade.

Re: GLM 5.2 beats Claude in our benchmarks

#48
post #42

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

I think state-of-the-art AI is going to be defense industry only from now on. We can have our toy drones but not the Predators and Reapers.

the things that empower modern toy drones were export restricted for years before hand.

Re: GLM 5.2 beats Claude in our benchmarks

#50
post #10

Earlier quoted context omitted.

I think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"

Are you replying to a response to the original comment? I looked but i didn't see anyone saying he's throwing shade.

You have to forgive the GLM bot. It's not very good.
Post reply on HN