Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

541–550 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#541
post #262

Earlier quoted context omitted.

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way mo…

Because it truly makes a difference. Opus 4.8 was great until we experienced Fable 5.

And post Fable retraction, I am now most certaily noticing Opus being 'dumber' also.

Open Weights are good. Not (yet) as good as leading closed models. Unfortunatly they will be declared 'illegal' any day now, and I unfortunately do not see myself able to run GML 5.2 in my basement homelab any time soon.

Re: GLM 5.2 beats Claude in our benchmarks

#542
This is an interesting finding, but very specialised. It would also be great to get some more information about the benchmark. Is it just a collection of files with vulnerabilities, or are they hidden in a real codebase, where LLM based approaches will not be able to scan every file like a static code scanner is able todo.

Re: GLM 5.2 beats Claude in our benchmarks

#544

Earlier quoted context omitted.

For me, the 20€/months subscriptions were always sufficient, and it's nice if that subscription give the latest and greatest results.

It depends. Claude’s $20 plan is kind of a mess.

That's exactly the one I'm really satisfied with.

Re: GLM 5.2 beats Claude in our benchmarks

#545

Earlier quoted context omitted.

We need a benchmark of independent community sourced benchmarks! …probably already is one

I don't know how you'd judge benchmarks beyond "did it test and measure what it says it tests and measures". And, I guess there have been instances where the benchmark failed to do that, and the models could cheat in some way and it just tested the models ability to find the answer key. In the case of my benchmarks every model other than Claude models running in Claude Code never have network access and all informati…

Almost glad I was dreadfully unclear b/c that’s interesting! What I had in mind, myself-

Thinking more a project to collect the results of benchmarks run by independent tech professionals and hobbyists, then compute a new meta-benchmark based on that real-world feedback.

Re: GLM 5.2 beats Claude in our benchmarks

#547

Earlier quoted context omitted.

Why not? Mythos level really doesn't seem that scary. And it would be a great way to take away the American labs international market. I think it would make strategic sense for them to release more capable models than what American labs are allowed to make available to the world. It would help them grow their global soft-power and be a destabilizing effect on the American economy.

It is fairly obvious to me that the open models are a form of "dumping" as far as the economics and the desired outcome from China's perspective. They get to watch as the US pours tons of money and talent into an industry, then prevent that investment from having any return. In 5 years we'll be on equal footing, China will have spent 1/1000th the money, and the only downside will be that they spent 5 years being 6 mo…

This assumes that AGI is not a winner takes all or winner takes most scenario. We just don't know that. If RSI actually happens, then it really could be winner takes all, and in that case being 6 months behind is the same as not even playing at all.

Re: GLM 5.2 beats Claude in our benchmarks

#548

Earlier quoted context omitted.

You are delusional if you think China is going to let Europe have access to Mythos level models for free.

Why not? Mythos level really doesn't seem that scary. And it would be a great way to take away the American labs international market. I think it would make strategic sense for them to release more capable models than what American labs are allowed to make available to the world. It would help them grow their global soft-power and be a destabilizing effect on the American economy.

>Mythos level really doesn't seem that scary.

This is not a serious comment. A model capable of chaining exploits together is capable of causing massive damage through cyber attacks.

Re: GLM 5.2 beats Claude in our benchmarks

#549

Earlier quoted context omitted.

You might feel different if you're a palestinian who's getting american bombs dropped on him, or an afghani collateral damage or... There is no good guys in general, and whataboutism and making the scope bigger doesn't help. The thing is that if the models you are building on are open source whether hosted on chinese / american / whatever service at least give you an option to switch provider easier vs a fable / chat…

> palestinian who's getting american bombs dropped on him So far as I know, there have not been any offensive American operations in the West Bank or Gaza. Do you have sources for any? > There is no good guys in general, and whataboutism and making the scope bigger doesn't help. The thing is that if the models you are building on are open source whether hosted on chinese / american / whatever service at least give yo…

1. I said American bombs. I didn't say dropped by American pilots. But provided, directly by the USA to Israel.

2. The whole idea is that open source means you don't have to care about "native". Look up nginx for a great example.

3. America from ww1 ish, definitely after ww2 and until about 2005/10 had a very very good reputation in most of Europe. The obama years gave it a bump up as well So basically for about 80-90% of the last 100yrs.

Re: GLM 5.2 beats Claude in our benchmarks

#550

Earlier quoted context omitted.

Is this 2006? Linux is present on literal billions of android phones, servers, supercomputers and other embedded devices. It's the most ubiquitous OS on the planet and it's not even close, even Microsoft contributes to it. The only niche where it doesn't utterly dwarf the competition is personal computers and it looks like we're all getting priced out of that anyway

I was 100% sure that somebody will throw this, but I didn't actually expect to do it from a throwaway account. Maybe because you know calling Android Linux is like calling a human just an ape (but mirrored because you know, Android in this case is the ape). Oh, and Microsoft Loves Linux, right? Because that's why they invented WSL, to make people go and use Linux, right? riight!!

Moving goalposts. Android is Linux (the GNU part of GNU/Linux is trivial to add), people use Linux knowingly or not, and Linux is on the most devices on Earth regardless of whether their users know it or not.
Post reply on HN