Live data from Hacker News

GLM-5.2 is the new leading open weights model on Artificial Analysis

artificialanalysis.ai

411–420 of 476 posts

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#411
post #297

Earlier quoted context omitted.

Neuralwatt ... When you reverse calculate the actual energy usage / price on a token basis, the gap is large. I do not have GLM 5.2 numbers because the whole default max setting is overkill. But GLM 5.1 numbers had it at 12x cheaper then API rates. And about 2.5x more tokens vs zai their own subscription service. Yes, its FP8 but lets be honest, do we know for sure that even zai runs at FP16? I learned a long time ag…

Please correct me if you have contradicting data but: Neuralwatt's price per token vs price for energy comparison doesn't seem to take into account the cost savings from cache hits that other providers offer on pure token rates. The comparison seems to assume every input token is a cache miss. On top of that, the cloud offering doesn't seem that well-run, they randomly blocked a colleague's API key for a couple days…

> doesn't seem to take into account the cost savings from cache hits

Absolute false information.

From my usage panel for this month:

* Total Tokens 1.1B * Cached Tokens 1.0B 97% of prompt tokens * Cost energy pricing $26.58

The energy pricing is higher then what i actually pay because its a mix of token billing and partial subscription (60% extra "power").

From the $50 subscription, i have about 3/4 left (4.21 of 16.0 kWh used this billing cycle). Used $5.5 in token billing.

That was running 82.0% GLM 5.1, and 18% GLM 5.2. Yes, i have been busy ;)

My actual usage if we look in dollar value was ~ $18.

For your information, that is cheaper the MiMo v2.5 Pro from Xiaomi as there i was doing around 450.000t per cent. And they have the same 75% cheaper prices like DeepSeek. MiMo has a issue with cache retention between session prompts what hurts them vs DeepSeek. Yes, DeepSeek v4 Pro is 2.5x cheaper but nowhere near GLM 5.1, and especially not GLM 5.2.

In case your wondering, zai subscription light is about 80m token / week limit. So on a token/cent price, neutralwatt is about 3x cheaper (and not 5h, week limits to maximize/frustrate).

> all while taking weeks to onboard new models.

Took them 1 day to include GLM 5.2 ... Yes, the remove old models fast because they do not have the server capacity to keep old models around.

> I assume some of these problems would be addressed if we had an SLA/enterprise contract.

Its a small team, not a big huge company. From my experience so far, seen a 2 timeouts, and sometimes slow speeds as servers get overloaded. For what i am paying for GLM ~5.1~ 5.2 ...

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#412
post #135
post #116

Earlier quoted context omitted.

I don't see this being such a big gap. There are some use-cases for sure but apart from UX/UI work it is not really needed. Besides, none of the frontier models can replicate actual images - the can approximate at least in my own experience.

One of my tests for a new model is dumping in a screenshot of a web page and seeing if it can recreate it from scratch in HTML and CSS. Even the local models I run on my Mac are getting surprisingly good at that now.

a pretty fun and quick tests i do with vision models is to screenshot the hackernews homepage and ask the model to return a json representation of the screenshot - qwen 3.5 0.8b did surprisingly well at this.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#413

GLM 5.2 is the first model we've tested that is unambiguously on par with, or better than Opus 4.6 (although as usual, we have GLM 5.2 and most other Chinese models a bit below most other benchmarks with more vulnerable test methodologies). Data at https://gertlabs.com/rankings

I really have to take your score with a grain of salt because Opus 4.5 does better than Opus 4.6

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#414
post #4

Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw) This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API p…

Which of those providers are:

1. Keeping your data private on in the US

2. Not training on it

3. Not quantizing the model

4. Offer reasonable latency adds rate limits

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#415
post #4

Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw) This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API p…

[flagged]

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#416
post #4

Why aren't more people talking about this? It's literally Opus 4.7 quality stupid prices. I know providers who are offering this at unlimited tokens for $50 a month. Some are even offering API rates at 3x lower than the official ZAI api rates which are already like 10x cheaper than Opus. (Crof and Umans btw) This is a huge blow to Anthropic/OpenAI/Google and a massive win for the rest of the world. The official API p…

Be careful about unofficial providers, a lot of them misconfigure models or stealth quantize them. For a while the difference between Kimi on the official API and most third party providers was 20-40%.

[flagged]

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#417

Earlier quoted context omitted.

This is a problem I find with opus is will spend so long thinking then going “but wait what if” To point where I stop it and simple tell it to “start writing code you can work it out as you go along” Seems writers block also effects LLM

https://arxiv.org/abs/2606.00206 In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small models in the paper) terminate reasoning faster and perform better. I bet Anthropic is tuning this on their backend.

Didn't they originally introduce those tokens to make the models smarter by second guessing their "thoughts"?

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#418

So this basically means we will have a near opus level model able to be run locally in the next couple of months right? QWEN 3.6 27b is already pretty good, but it should be possible to get a better option now that runs in the same hardware, right?

So much depends on the thinking effort, it's almost meaningless to compare these models without specifying it. GLM 5.2 needs to run with max thinking effort to be competitive with the leading-edge models from OpenAI and Anthropic. That slows it down quite a bit in my experience. Meanwhile, those models have thinking-effort knobs of their own that make a big difference, especially in GPT 5.5's case.

I have been messing with an early NV4FP quant of GLM 5.2 and so far, that model in its Max setting outperforms GPT 5.5 on its default setting. But GPT 5.5 still pulls ahead once I crank up its own reasoning effort. I imagine the same is true of Opus 4.x but haven't pitted them against each other yet.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#419
post #110

I was surprised that GLM 5.1/5.2 are not vision models - they are text input only. That's actually pretty uncommon these days. All of the OpenAI/Anthropic/Gemini models accept images, and so do the other leading open weight families - Gemma 4, Qwen 3.6, Kimi 2.x. In GLM's case image input would be useful because it's a model that scores very highly for tasks like web design, but without image input it can't take a sc…

Agreed, that's actually one step that will make people adopt it widely for customer facing AI Agent!

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#420
post #297

Earlier quoted context omitted.

Please correct me if you have contradicting data but: Neuralwatt's price per token vs price for energy comparison doesn't seem to take into account the cost savings from cache hits that other providers offer on pure token rates. The comparison seems to assume every input token is a cache miss. On top of that, the cloud offering doesn't seem that well-run, they randomly blocked a colleague's API key for a couple days…

> doesn't seem to take into account the cost savings from cache hits Absolute false information. From my usage panel for this month: * Total Tokens 1.1B * Cached Tokens 1.0B 97% of prompt tokens * Cost energy pricing $26.58 The energy pricing is higher then what i actually pay because its a mix of token billing and partial subscription (60% extra "power"). From the $50 subscription, i have about 3/4 left (4.21 of 16.…

Your reply doesn't seem to be in good faith. Please provide your formula for calculating effective per token cost.

I am not sure why the small team argument is relevant. This is a crowded market, there are dozens if hundreds of third party inference providers in the world right now. I'm glad that's a good excuse that works on you but I'm not sure why the average user should care.

Post reply on HN