Live data from Hacker News

GLM-5.2 is the new leading open weights model on Artificial Analysis

artificialanalysis.ai

351–360 of 476 posts

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#351
post #71

Artificial Analysis coding benchmark shows GLM5.1 on high pretty close to GPT5.5 xhigh in cost to run, with GPT5.5 on medium significantly less expensive. Compared to GPT5.5 medium GLM5.1xhigh is twice the cost and half the intelligence. They don't have GLM5.2 on there yet, but that'd a big gap to bridge. https://artificialanalysis.ai/agents/coding-agents?coding-ag... I thought I was "holding it wrong" until DeepSWE…

It got 46.2 on DeepSWE in Z.ai's own run[1]. That would put it between Opus 4.7 xhigh and Opus 4.8 medium. [1] https://z.ai/blog/glm-5.2

If that ends up being true, GPT5.5 at 70 (and presumably Fable a bit ahead of that) is still in a different league, which was partly my point. To listen to online chatter, GLM5.2 is a tectonic shift in the landscape. In reality, it's just interesting. Probably safe to bet once the DeepSWE benches all get fully updated it won't even be on the pareto frontier.

I'm not accusing anyone specifically, but I've noticed Chinese bots swamping certain YouTube channels that, for example, cover US defense industry news. They'll downplay any and all technical advances, play up China's dominance, US cowardice, etc. All very transparent. I suspect some of the online conversation about open Chinese models is driven by that. How often do you see people talking about Mistral or Trinity? Never. Because they don't play that game.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#352

Earlier quoted context omitted.

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

For Claude models at least, you can tell to just manually think in the output and it works fine. I do it reguralrly because for creative writing and summarization, they seem to believe they don't need to think at all, and get way worse results.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#353
post #350

Earlier quoted context omitted.

Sir, I would suggest that if Europe fails to be economically competitive, the downstream implications on European society will produce much worse outcomes than (for instance) data transparency… Doing things with ethical intentions does not necessarily produce outcomes that are beneficial for society at large.

Well, is this mad dash for AI producing "outcomes that are beneficial for society at large" yet? So far it looks like its mostly producing a ton of negative externalities and wealth transfer to corrupt elites. Also, no, abandoning ethics is not an option, what a ridiculous suggestion.

Data transparency and copyright does not constitute “ethics.”

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#354

Earlier quoted context omitted.

with open models you can get a subscription with privacy, at the same cost as codex. openai, google and anthropic subscriptions are not available with privacy. looking at the link there it's interesting that going from cursor cli to codex cli take gpt 5.5 from 7th to 3rd. but they didn't do open model in codex. so, hard to say it's for sure a model benchmark. maybe open models are just shit at swe agent harness...it'…

> with open models you can get a subscription with privacy Unless you're running it locally, aren't you just trusting some other entity?

right, and on prem being an option is a god send, however you manage to do it

it's not a recommendation, its an option. if you don't have capital then it doesn't apply to you and move on. it wasn't an option for even people with capital.

come back in a few years when its more accessible

additionally I like that there are providers with faster special purpose processors for faster tokens/sec, all at different pricing strategies

so just pick something that matches your personal risk tolerance

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#355
post #108

Earlier quoted context omitted.

IME, unquantised -> FP8 is pretty much lossless. What matters more is having an unquantized KV cache - using an FP8 KV cache can result in a significant drop in quality.

>unquantised -> FP8 is pretty much lossless Claude Shannon is rolling in his grave.

I don't know, sounds quite similar to his rate distortion theorem (analyzing minimum number of bits/symbol you need to stay under some fixed amount of distortion). I.e. lossy compression with a maximum amount of loss. I.e. "pretty much lossless" compression.

https://en.wikipedia.org/wiki/Rate%E2%80%93distortion_theory

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#356

GLM 5.2 is the first model we've tested that is unambiguously on par with, or better than Opus 4.6 (although as usual, we have GLM 5.2 and most other Chinese models a bit below most other benchmarks with more vulnerable test methodologies). Data at https://gertlabs.com/rankings

[dead]

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#357
post #4

Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw) This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API p…

To answer the question in your first sentence - because it's VERY computationally (ha) expensive as a human being to keep up with all the options. It's also very hard to figure out how to run a model like this. There's no installer . If you really really care, which 99% of people do not, you have to google a guide, and then find out it's out of date... I've tried a number of these, and the learning curve is very stee…

It's also very hard to figure out how to run a model like this. There's no installer.

Yes, there is. It's called Claude Code. Point it at the HuggingFace URL and say "Download these weights and build whatever is needed to run them, then test the model."

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#358
post #351

Earlier quoted context omitted.

It got 46.2 on DeepSWE in Z.ai's own run[1]. That would put it between Opus 4.7 xhigh and Opus 4.8 medium. [1] https://z.ai/blog/glm-5.2

If that ends up being true, GPT5.5 at 70 (and presumably Fable a bit ahead of that) is still in a different league, which was partly my point. To listen to online chatter, GLM5.2 is a tectonic shift in the landscape. In reality, it's just interesting. Probably safe to bet once the DeepSWE benches all get fully updated it won't even be on the pareto frontier. I'm not accusing anyone specifically, but I've noticed Chin…

There are definitely some Chinese bots + actual people (imagine that!) who like to talk up Chinese models, I'm one of them but I like to find out how good these models really are before saying anything.

GLM definitely isn't opus level yet but it's for sure good. I think it lacks some knowledge (when coding) that the frontier models possess, which is expected given that the model is probably quite small when compared to the frontier.

But people don't say much about Mistral, probably because they are nowhere as good.. And they don't have large population behind them to actually use them.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#359
post #2

It seems to really be a nice step-up and is getting quite close to the frontier. I wish they'd start focusing on the reasoning efficiency now, though. I have a simple (relatively) test task to evaluate LLMs: writing a simple math evaluator library in Nim (it's about 400-600 lines total max), and GLM 5.2 (xhigh which maps to max effort) spent over 15 minutes (!) reasoning, spending about 45k tokens, before it finally…

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

> GLM 5.2 Max = Opus 4.8 Max in thinking behavior

This is insane! I can't wait until technology progresses to the point we can run these things on consumer hardware!

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#360

I have a script that ranks these based on codingindex from Artificial Analysis. All it does is pull a json from their main table page and parses it with the fields I care about (coding). There used to be a mailing list associated with it but eh ... there wasn't much interest. I use the script every day though. Current partial output score age size name 47.1 58 large Kimi K2.6 47.5 54 large DeepSeek V4 Pro (Reasoning,…

[dead]
Post reply on HN