Live data from Hacker News

GLM-5.2 is the new leading open weights model on Artificial Analysis

artificialanalysis.ai

221–230 of 476 posts

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#221
post #183

The problem with these benchmarks is that the Chinese models tend to be incredible on paper, and absolutely terrible in practice :/

I beg to differ. I replaced a $40/mo GitHub Copilot subscription where I used Opus 4.6 and GPT 5.5 with a $10/mo opencode Go plan where I use mostly DeepSeek V4 Flash and testing MiMo 2.5. I work on mid-sized projects currently (200k to 1kk lines of code).

> 1kk lines of code

Isn't that a million?

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#222
post #69

Earlier quoted context omitted.

they went public a few weeks ago

That's cool and all, but they are still on GLM 4.7

Which is fine for their target market. Their latest model is Kimi K2.6, available to enterprise customers. But older models become more powerful when you have time to do more reasoning. Also many applications don't need advanced models. Cerebras is making bank from all the other use cases that SOTA providers left on the table by focusing on 0-shot intelligence over speed

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#223

The problem with these benchmarks is that the Chinese models tend to be incredible on paper, and absolutely terrible in practice :/

You are obviously lying because it shows you have no experience with. GLM since 4.5 have been crushing it. all their models since then haven't skipped a beat. 4.5/4.5-air, 4.6, 4.7, 4.8, 5, 5.1. That aside, MiMoV2.5, MiniMax from 2.0, DeepSeek from V3, Kimi since V2, Qwen since 3, Hy3 have all been amazing models. All from China, we need to get over it. China is not losing yet as far as the AI race is concerned.

Is there a GLM-4.8 model?

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#224

The problem with these benchmarks is that the Chinese models tend to be incredible on paper, and absolutely terrible in practice :/

I have used GLM since version 4.8 I think and do enjoy using them. More then other models like Kimi or Deepseek. Though only tested them on smaller private projects.

> I have used GLM since version 4.8 I think

You probably refer to GLM-4.7

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#225

Earlier quoted context omitted.

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

>such outrageous copyright infringement Sarcasm, considering the source of their own training data?

IP for me, not thee.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#226

Knowing very little about how to run these, how close are we to medium or larger businesses starting to buy hardware to run models like this to keep the models local? It’s expensive, and not as capable as the frontier models, but would have some pretty big benefits around privacy and agency.

I know of multiple businesses in Europe that have been doing that for a while with 70B models, and are upgrading hardware to run the new crop of 700B-1T models (really started around Kimi K2, but buying and hosting that kind of hardware takes time) Not everyone is willing (or even legally able) to send their trade secrets to OpenAI or Anthropic

While certainly there are such cases with trade secrets, it's worth noting that even large banks typically have a provider like Azure or AWS onboarded.

There they can deploy these models while using the existing legal frameworks.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#227

I have a script that ranks these based on codingindex from Artificial Analysis. All it does is pull a json from their main table page and parses it with the fields I care about (coding). There used to be a mailing list associated with it but eh ... there wasn't much interest. I use the script every day though. Current partial output score age size name 47.1 58 large Kimi K2.6 47.5 54 large DeepSeek V4 Pro (Reasoning,…

score age size name 62.0 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.1 55 - GPT-5.5 (xhigh) 58.5 55 - GPT-5.5 (high) 57.2 104 - GPT-5.4 (xhigh) 56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 56.2 55 - GPT-5.5 (medium) 55.5 118 - Gemini 3.1 Pro Preview 53.1 132 - GPT-5.3 Codex (xhigh) 53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort) 52.5 62 - Claude Opus 4.7 (Adaptive Re…

[deleted]

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#229

Earlier quoted context omitted.

score age size name 62.0 8 - Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 59.1 55 - GPT-5.5 (xhigh) 58.5 55 - GPT-5.5 (high) 57.2 104 - GPT-5.4 (xhigh) 56.7 20 - Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 56.2 55 - GPT-5.5 (medium) 55.5 118 - Gemini 3.1 Pro Preview 53.1 132 - GPT-5.3 Codex (xhigh) 53.1 62 - Claude Opus 4.7 (Non-reasoning, High Effort) 52.5 62 - Claude Opus 4.7 (Adaptive Re…

Short comments... - GPT 5.5 consistently the best, an opinion who gets me constant downvotes here by the Anthropic Marketeer strike force... - China is going to eat the US lunch on AI - What have European universities and companies been doing? Its like if, on a parallel past/future, Nikola Tesla and Edison would have created flying Cyberpunk machines, while Europeans researchers, would be getting together to request…

"…Anthropic Marketeer strike force…"

Might also just be the result of "good will" (that the company has deftly fostered). Other companies might learn from Anthropic in that regard.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#230
post #71

Artificial Analysis coding benchmark shows GLM5.1 on high pretty close to GPT5.5 xhigh in cost to run, with GPT5.5 on medium significantly less expensive. Compared to GPT5.5 medium GLM5.1xhigh is twice the cost and half the intelligence. They don't have GLM5.2 on there yet, but that'd a big gap to bridge. https://artificialanalysis.ai/agents/coding-agents?coding-ag... I thought I was "holding it wrong" until DeepSWE…

with open models you can get a subscription with privacy, at the same cost as codex. openai, google and anthropic subscriptions are not available with privacy. looking at the link there it's interesting that going from cursor cli to codex cli take gpt 5.5 from 7th to 3rd. but they didn't do open model in codex. so, hard to say it's for sure a model benchmark. maybe open models are just shit at swe agent harness...it'…

> with open models you can get a subscription with privacy

Unless you're running it locally, aren't you just trusting some other entity?

Post reply on HN