Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

191–200 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#191

I added GLM 5.2 to my security bug hunting benchmark when it came out, and found it to be a good performer, but not the best open model. The benchmark tests whether models can find bugs Mythos found. The best open models in the initial benchmark were DeepSeek V4 Pro or MiMo 2.5 Pro. But it turned out MiMo got lucky, it's performed worse on almost every test I've done since, while DeepSeek has consistently been among…

We need a benchmark of independent community sourced benchmarks!

…probably already is one

Re: GLM 5.2 beats Claude in our benchmarks

#192
post #55

Earlier quoted context omitted.

8 X RTX6000. It will run you around 80-100k to get started with a model at this size with decent tps.. Don't worry though, open source evangelists will tell you that these will be running on your phone in the next 3 years. For $100k you could run this model 24/7 through open router with 10 concurrent sessions at 50tps for a decade and have money left over for a vacation. There's no point in investing this type of mon…

> 8 X RTX6000. It will run you around 80-100k to get started 8 x RTX6000 GPUs cost $100,000 alone. You then need to build a system that can support those GPUs with enough PCIe lanes through a PCIe switch. It's going to be $120K to $150K to build or buy a system to run this.

Not to mention the three separate dedicated 15A circuits you would need to have installed in order to run the 3x 2000W power supplies running ideally at no more than 1400W sustained load each. And definitely would need 200A service to the house if you have a family living there with you.

But hey you could save on heating?

Re: GLM 5.2 beats Claude in our benchmarks

#193

Earlier quoted context omitted.

If that happens it'll be an absolute disaster. Imagine a scenario where Anthropic and OpenAI prohibit most US companies from using their latest models because of safety.. And meanwhile attackers use equivalent open source models to attack US companies. Any prohibition on open source models will do nothing to fix the problem.. since attackers will never feel bound to the law. All advanced models must be available for…

It'd be less about "safety" and more "we've spent trillions developing these AI tools only to have the Chinese, once again, copy them and offer them for pennies on the dollar, and no one seems to care about the impact that has on the long-term sustainability of this sector of the American economy as a whole, so we're yanking the models."

"I'm going to take this box razor and make some really deep cuts around the middle of my face because my tech sector is too good and that's actually a bad thing because $foreigners."

Re: GLM 5.2 beats Claude in our benchmarks

#195

Earlier quoted context omitted.

Yea as far has hobbies go, I feel like this is on the low end. I know people who collect watches and corvettes, that's way more expensive and functionally you can't really do anything special with them.

The difference is watches and corvettes typically appreciate in value, where as computer hardware typically drops like a rock.

> watches

Some, and the market fluctuates a ton.

> corvettes

Only the oldest, most unique model years: nobody is buying (C4-C5-realistically C6) mid-90s or early 2000s Corvettes for more than what they paid for them, and they never will.

Re: GLM 5.2 beats Claude in our benchmarks

#196

I like GLM 5.2... ish. It's ok. I'd be mostly fine switching to it. I just can't find a cost effective way to do that. z.AI's coding plan is both overpriced and unreliable. ollama's is also overpriced. Paying by the token for it on openrouter etc is more expensive than just having a Codex or Claude coding plan. If you have to pay by the token, it's clearly cheaper. It's not competitive with a coding plan though.

It also means giving up vision which I don't know how I would deal with. I think I would prefer a weaker model with vision than a stronger without.

Openrouter definitely supports vision models. Why would you have to give up vision?

Re: GLM 5.2 beats Claude in our benchmarks

#197

Earlier quoted context omitted.

It will almost for sure surpass the models which Trump will allow US "allies" (which he just considers client states) to use. This, together with China's growing dominance in PV, rechargeable batteries, EV, could really be the nail in the coffin for the post WWII economic world order.

You are delusional if you think China is going to let Europe have access to Mythos level models for free.

Why not?

Mythos level really doesn't seem that scary. And it would be a great way to take away the American labs international market.

I think it would make strategic sense for them to release more capable models than what American labs are allowed to make available to the world. It would help them grow their global soft-power and be a destabilizing effect on the American economy.

Re: GLM 5.2 beats Claude in our benchmarks

#198

Earlier quoted context omitted.

Yea as far has hobbies go, I feel like this is on the low end. I know people who collect watches and corvettes, that's way more expensive and functionally you can't really do anything special with them.

The difference is watches and corvettes typically appreciate in value, where as computer hardware typically drops like a rock.

> The difference is watches and corvettes typically appreciate in value

Both of those things' value drops like a rock as soon as you buy them and, at least for cars, they don't all appreciate. Most don't. Even so, they appreciate at an incredible slow rate.

I can't speak for watches but I'd be surprised if it wasn't the same situation.

At least the gpus can create value after you buy them before they are worthless.

Re: GLM 5.2 beats Claude in our benchmarks

#199

These numbers are seem pretty low compared to what I was able to achieve specifically around windows kernel, win32k win32u to be exact. It honestly wouldn't surprise me anymore if china started surpassing models that US makes public, at least in specific categories such as cyber. GLM 5.2 is already capable enough to assist in self-training which is similar to what we saw happen with frontier models and they appear to…

> These numbers are seem pretty low compared to what I was able to achieve specifically around windows kernel, win32kwin32u to be exact.

Care to give more context to this? Seems interesting

Re: GLM 5.2 beats Claude in our benchmarks

#200

I added GLM 5.2 to my security bug hunting benchmark when it came out, and found it to be a good performer, but not the best open model. The benchmark tests whether models can find bugs Mythos found. The best open models in the initial benchmark were DeepSeek V4 Pro or MiMo 2.5 Pro. But it turned out MiMo got lucky, it's performed worse on almost every test I've done since, while DeepSeek has consistently been among…

We need a benchmark of independent community sourced benchmarks! …probably already is one

I don't know how you'd judge benchmarks beyond "did it test and measure what it says it tests and measures". And, I guess there have been instances where the benchmark failed to do that, and the models could cheat in some way and it just tested the models ability to find the answer key. In the case of my benchmarks every model other than Claude models running in Claude Code never have network access and all information from after the bug was discovered has been removed from the repository the model can see.

But, there are benchmarks for so many different kinds of ability, I don't know how to compare them directly against one another. Like, models that do well on terminal and agentic coding benchmarks tend to do well on finding security bugs, but it's not a 1:1 correlation, there are surprises.

Post reply on HN