Live data from Hacker News

GLM-5.2 is the new leading open weights model on Artificial Analysis

artificialanalysis.ai

361–370 of 476 posts

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#361

Earlier quoted context omitted.

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

I don’t understand why there isn’t public dataset for reasoning that can be improved by humans/llms like Wikipedia (ie with auto judging contributions etc).

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#362

Earlier quoted context omitted.

To answer the question in your first sentence - because it's VERY computationally (ha) expensive as a human being to keep up with all the options. It's also very hard to figure out how to run a model like this. There's no installer . If you really really care, which 99% of people do not, you have to google a guide, and then find out it's out of date... I've tried a number of these, and the learning curve is very stee…

It's also very hard to figure out how to run a model like this. There's no installer. Yes, there is. It's called Claude Code. Point it at the HuggingFace URL and say "Download these weights and build whatever is needed to run them, then test the model."

I really miss the time when people thought that the idea of someone telling an un-sandboxed AI "do whatever is needed to X" was unrealistically stupid.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#363

Earlier quoted context omitted.

score age size name 62.0 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.1 55 - GPT-5.5 (xhigh) 58.5 55 - GPT-5.5 (high) 57.2 104 - GPT-5.4 (xhigh) 56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 56.2 55 - GPT-5.5 (medium) 55.5 118 - Gemini 3.1 Pro Preview 53.1 132 - GPT-5.3 Codex (xhigh) 53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort) 52.5 62 - Claude Opus 4.7 (Adaptive Re…

rank score age size name 1 62.0 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 2 59.1 55 - GPT-5.5 (xhigh) 3 58.5 55 - GPT-5.5 (high) 4 57.2 104 - GPT-5.4 (xhigh) 5 56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 6 55.5 118 - Gemini 3.1 Pro Preview 7 53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort) 8 53.1 132 - GPT-5.3 Codex (xhigh) 9 52.5 62 - Claude Opus 4.7 (Adaptive Reason…

These results are amazing! I can't believe an open weight model rivals Opus 4.6, my most used model!

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#364
post #352

Earlier quoted context omitted.

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

For Claude models at least, you can tell to just manually think in the output and it works fine. I do it reguralrly because for creative writing and summarization, they seem to believe they don't need to think at all, and get way worse results.

this helps so much. i do it too. with some of the newer frontier models its unclear if you can even turn it off in the first party chat apps. havent compared api semantics yet.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#365
post #110

I was surprised that GLM 5.1/5.2 are not vision models - they are text input only. That's actually pretty uncommon these days. All of the OpenAI/Anthropic/Gemini models accept images, and so do the other leading open weight families - Gemma 4, Qwen 3.6, Kimi 2.x. In GLM's case image input would be useful because it's a model that scores very highly for tasks like web design, but without image input it can't take a sc…

Configure a subagent in your coding harness to spin up a new sub-session with any vision model for those tasks and feed the result back to the main model. No need for "one model that does everything"

That doesn’t work well in a lot of scenarios. The text LLM doesn’t know what to look for in an image before it sees a description, you might need multiple rounds of back and forth.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#366

Earlier quoted context omitted.

It's also very hard to figure out how to run a model like this. There's no installer. Yes, there is. It's called Claude Code. Point it at the HuggingFace URL and say "Download these weights and build whatever is needed to run them, then test the model."

I really miss the time when people thought that the idea of someone telling an un-sandboxed AI "do whatever is needed to X" was unrealistically stupid.

Skill issue

(In all seriousness, I agree this is a problem. That capability is too powerful not to take advantage of, though. Nobody needs to struggle with this sort of thing anymore, but yes, obviously, it should happen in a VM or at least a container.)

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#367
Also so wild that it's relatively compact. 753B-40A is so reasonable, shows incredible scaling in what the model can do, without just throwing heaps of new parameters in.

This is silly but I dig how 753 is very close to 745, which is the watts in a HP. 1bHP parameter model. Silly, but I enjoy it.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#368
Seems really good at frontend work, and as a result on remotion programmatic videos. Not the best yet, thats still Gemini 3.1 pro(trained on actual videos) or Fable, but often better than what Opus can come up with

https://mesmer.tools/benchmarks/ai-video-generation

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#369
post #4

Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw) This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API p…

Isn't it closer to sonnet?

The Chinese open weight models have been ahead of Sonnet (at least for coding) for a couple months now. I tend to take benchmarks with a huge grain of salt, but in my own experience, the latest versions of Kimi, MiMo, and GLM (pre-5.2) had already surpassed Sonnet in terms of output quality for a fraction of the price.

With that said, I'm excited to try GLM 5.2 because I still end up reaching for Opus and GPT 5.5 for many tasks because the open models tend to get stuck more often on complex problems.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#370

Earlier quoted context omitted.

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

You can trivially leak the CoT of any current model, it's not a problem.

>outrageous copyright infringement

>unethically scrapped data

Hahahahaha

Post reply on HN