Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

521–530 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#521

Earlier quoted context omitted.

Yea as far has hobbies go, I feel like this is on the low end. I know people who collect watches and corvettes, that's way more expensive and functionally you can't really do anything special with them.

The difference is watches and corvettes typically appreciate in value, where as computer hardware typically drops like a rock.

Corvettes don't appreciate in value, and high end data center hardware isn't dropping in value anymore. A100s are more than 2 dollars an hour, more than they cost in 2023.

Re: GLM 5.2 beats Claude in our benchmarks

#522
post #451

I hope someone is also building a Claude Design competitor. One that is similarly HTML based instead of the Figma/Magic Patterns approach. I have more vendor lock-in with Design than I do with Code, and will switch over as soon as Claude loses the smallest technical advantage

https://github.com/nexu-io/open-design

Re: GLM 5.2 beats Claude in our benchmarks

#523
post #262

Earlier quoted context omitted.

> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way mo…

For me, the 20€/months subscriptions were always sufficient, and it's nice if that subscription give the latest and greatest results.

It depends. Claude’s $20 plan is kind of a mess.

Re: GLM 5.2 beats Claude in our benchmarks

#524

Earlier quoted context omitted.

x40hrs/week * 50 weeks = $180k Congrats, now you’re paying an engineer’s salary to make your engineer at best 20% more productive. Better to hire another engineer, or two jrs, and build up your in house talent.

Only you get things done lost faster and don't need to pay entire years salary?

How so?

If you’re doing a one-off project, sure. But if you’re coding like this full time, you’re paying the year’s salary anyway.

And faster? Not so sure about that. Sure they can write code faster, but writing code is a small part of building something.

Re: GLM 5.2 beats Claude in our benchmarks

#525
post #440

Earlier quoted context omitted.

x40hrs/week * 50 weeks = $180k Congrats, now you’re paying an engineer’s salary to make your engineer at best 20% more productive. Better to hire another engineer, or two jrs, and build up your in house talent.

except this is way more than an engineering salary. At least in Europe.

I’m sure there are engineers making $180k usd / year in the eu. Maybe it’s unusual, but hey, now you can cancel your claude subscription and hire a really good engineer

Re: GLM 5.2 beats Claude in our benchmarks

#526
post #413
post #408

Earlier quoted context omitted.

I would say 3.5 flash is great if you use a good open harness. I use omp for that. The thing with Google is that they announce they have a great model, and that they have been testing it internally for half a year. I guess they don't care too much about who or how he uses it. I am still struggling how to deal with sub agents and different roles for each model. I still think Claude or Codex are overall better models,…

Big fan of Amp but pretty sure it only uses Flash for search: https://ampcode.com/models As for Fable: I used it as much as I could while we had it. It was a step change over Opus with my work.

Or maybe it was supposed to be the OMP (OhMy Pi) harness. Pi can do just about anything for you. Use most models in most ways possible. You just tell Pi what you want, and it builds an extension for itself.

Re: GLM 5.2 beats Claude in our benchmarks

#527

Earlier quoted context omitted.

You might feel differently if you were a Filipino or Vietnamese fisherman whose family relied on the income from the stocks of the South China Sea, or a Uighur person living in Western China, or a Ukrainian soldier who has to deal with drones built with Chinese components, or a democracy advocate in Hong Kong, or arguably, a person who had plans for 2020-2021. Or, on a more local note, an Australian automotive worker…

You might feel different if you're a palestinian who's getting american bombs dropped on him, or an afghani collateral damage or... There is no good guys in general, and whataboutism and making the scope bigger doesn't help. The thing is that if the models you are building on are open source whether hosted on chinese / american / whatever service at least give you an option to switch provider easier vs a fable / chat…

> palestinian who's getting american bombs dropped on him

So far as I know, there have not been any offensive American operations in the West Bank or Gaza. Do you have sources for any?

> There is no good guys in general, and whataboutism and making the scope bigger doesn't help. The thing is that if the models you are building on are open source whether hosted on chinese / american / whatever service at least give you an option to switch provider easier vs a fable / chatgpt 5.6 that gets banned for none americans etc...

Then it'd make far more sense to use a native open-source AI project than to use Chinese AI. Personally, I've been looking into Mistral AI for my own uses.

> 2 years ago america would have had the branding/perception advantage but right now that is well and truly gone...

Hardly. Let's not pretend that much of the world, particularly Europe, hasn't had a long and storied history of staring down its nose at Americans. At a certain point you just stop caring, regardless of your political persuasion.

Re: GLM 5.2 beats Claude in our benchmarks

#528
post #262

Earlier quoted context omitted.

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way mo…

I am forced to use AI as part of my job to write code. As a matter of fact, I was recently told that I'm not using enough AI according to their metrics, even though I'm producing good quality code on time. Since the cost is one of the things I'm being judged on, you're damn right I want to use the newest and most expensive model available.

Re: GLM 5.2 beats Claude in our benchmarks

#529
post #59

I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and Open…

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

Something I don't see in your charts is acknowledgement of the difference, sometimes paradoxical, in strength between the same model at different reasoning levels. Do you have charts that include low/med/high/xhigh/max for the various models?

Re: GLM 5.2 beats Claude in our benchmarks

#530
I've been using it for a week via opencode in a large, mature codebase for some moderately ambitious feature development, and a bit of debugging. Explicit purpose is evaluating if it may be a good substitute to save money for many tasks. For several tasks I've had both it and opus 4.8 attempt the same task and compared them.

In general, it's comparable across the board. Claude is less "verbose" -- GLM really likes to comment a ton. There were a few things where I think claude would have needed a little bit less back and forth. So opus still has an edge, but it's marginal, very much unlike previous open/competitor models where benchmarks looked good but actual day to day performance was pretty bad. I'm sure fable is "better" but it's so expensive + data retention policies are such that for the moment it was generally available I couldn't use it for work. This is still notably better performance than when claude code took the industry by storm.

I'm understanding why Dario is trying to regulate open weight models away.

Post reply on HN