Live data from Hacker News

GLM-5.2 is the new leading open weights model on Artificial Analysis

artificialanalysis.ai

391–400 of 476 posts

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#391

Earlier quoted context omitted.

https://arxiv.org/abs/2606.00206 In this paper they nerf an LLMs ability to emit waffling thinking tokens like "wait", "but", "alternatively", and the models (they're old, small models in the paper) terminate reasoning faster and perform better. I bet Anthropic is tuning this on their backend.

This is super cool. Do you know if any of the inference backends (llama.cpp, vllm, etc) support this technique?

vLLM supports "banning" certain tokens but I don't know if it can dynamically reduce them.

To my knowledge you can also "ban" with llama.cpp but it is passed in the API call rather than to the server at initialization.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#392

Earlier quoted context omitted.

Apparently because of how Claude is trained, even the system level prompts go through as XML, it works better with XML "prompting" so I figured I could have it write plans in XML. I need to update my ticketing tool to output XML maybe by default. https://www.reddit.com/r/ClaudeAI/comments/1psxuv7/anthropic...

Comments later in thread say markdown works just as fine and that it’s more important to organize your plan into sections. Also just think about it, why would a model trained on the world’s corpus of text (that isnt formatted in xml) perform better with XML? It would be a better study if that post tested markdown, org, xml, json, etc. 10 times to see if their is a difference

XML consistently performed better than markdown and JSON in all evals I've ever seen on any model, except for a couple very specific ones.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#393

Earlier quoted context omitted.

Short comments... - GPT 5.5 consistently the best, an opinion who gets me constant downvotes here by the Anthropic Marketeer strike force... - China is going to eat the US lunch on AI - What have European universities and companies been doing? Its like if, on a parallel past/future, Nikola Tesla and Edison would have created flying Cyberpunk machines, while Europeans researchers, would be getting together to request…

> - If Zuckerberg could be fired, after spending a total of $235 billion on AI and having NOTHING to show for...should he be fired? Yes, if the premise was true but it’s not. https://opper.ai/ai-roundtable/questions/bbf5a4e9-204

Interesting...but this shows how dumb these AI are.

And they misunderstood nothing to show for as...literally nothing to show for. Yes not factually but he has nothing effectively not much that is competitive to show for so its literally true.

And had they been give this clarification then would have suddenly said: "Oh yes of course, you are absolutely right, you are correct on challenging me on that...."

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#395

Earlier quoted context omitted.

Isn't it closer to sonnet?

The Chinese open weight models have been ahead of Sonnet (at least for coding) for a couple months now. I tend to take benchmarks with a huge grain of salt, but in my own experience, the latest versions of Kimi, MiMo, and GLM (pre-5.2) had already surpassed Sonnet in terms of output quality for a fraction of the price. With that said, I'm excited to try GLM 5.2 because I still end up reaching for Opus and GPT 5.5 for…

I found sonnet preferable to k2.6 but 2.7 code for kimi seems better anecdotally

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#396

Earlier quoted context omitted.

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

distillation of thinking models is not particularly effective - both "Open"AI and Misanthropic don't show you the real chain of thought, only its severely downscaled version. both do everything in their power to combat such outrageous copyright infringement, so the bulk of unethically scrapped data the Chinese have is from several generations ago.

Reasoning models can coaxed to reason like they do in dedicated reasoning blocks, outside of those blocks: in normal parts of the response.

But Anthropic at least has openly admitted they try to detect that and interfere

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#397
post #81
post #53

In my tests[0] GLM-5.2 is not much better than GLM-5, and overall DeepSeek V4 Flash seems to be the better/more cost-effective choice: [0]: https://aibenchy.com/compare/deepseek-deepseek-v4-flash-high...

I think the problem is, as can also be seen on other benchmarks, is that most models nowadays are focused more and more purely on tool calling and coding. This means, that models are losing more and more general and domain-specific knowledge. Look at those graphs on ARtificialAnalysis, GLM-5.1 still performs similarly or better: AA-Omnisicence Accuracy: https://i.snipboard.io/5DYmpx.jpg IFBench: https://i.snipboard.i…

Well, in that example it still seems the big players are increasing overall "intelligence" as Fable tops the list.

OpenAI has big incentives to improve general interligence as a large percentage of users use ChatGPT for support, finances, questions, etc. Not just coding.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#398

Earlier quoted context omitted.

> GLM 5.2 Max = Opus 4.8 Max in thinking behavior This is insane! I can't wait until technology progresses to the point we can run these things on consumer hardware!

Are there any indications that this will be possible? Consumer hardware will continue getting better but I can't see 512GB RAM in a MacBook Pro any time soon. I'm hoping linear attention techniques plus MoE will make breakthroughs in size/compression and throughput.

Well, we're probably not going to be running frontier models anytime soon, but I think the general assumption is smaller models will continue to improve until they're sufficiently good frontier models aren't needed.

There's potentially also augmentation through tools, harnesses and RAG to help boost how well they work without tons of parameters.

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#399

Earlier quoted context omitted.

GLM 5.2 Max = Opus 4.8 Max in thinking behavior. The thinking chain is so similar, and so is the amount of token usage on the output. If you want reasonable token usage, you need to run it GLM 5.2 at High. There is little drop in quality from Max to High (for most tasks). And it cuts token usage by 2 a 2.5x. GLM 5.2, Max is really something you only need for complex tasks. In essence, GLM 5.2 is Opus 4.8 its little b…

> GLM 5.2 Max = Opus 4.8 Max in thinking behavior This is insane! I can't wait until technology progresses to the point we can run these things on consumer hardware!

you need 8 x 96GB Blackwell or equivalent

so around US$150k which is Small/Medium-Enterprise territory already, but who knows when it will hit "reasonable" home consumer territory

I think there's hope future generations of unified memory machines may get this sort of memory availability when new fabs open in then next couple of years and then ramp up production for a few years afterwards - that makes ~2030s credible at this point, but nobody can really predict the market that far ahead

Re: GLM-5.2 is the new leading open weights model on Artificial Analysis

#400
post #108
post #26

Earlier quoted context omitted.

> Some are even offering API rates at 3x lower than the official ZAI api rates Looking at openrouter [1], some of the cheaper offerings are for quantized models. Not sure how much intelligence is lost in quantization. And they are not 3 times cheaper. Where did you find 3x lower prices for APIs? I am considering skipping open router and using them directly for that price. edit: I see, croft [2] 8bit for $0.50/$0.08/$…

IME, unquantised -> FP8 is pretty much lossless. What matters more is having an unquantized KV cache - using an FP8 KV cache can result in a significant drop in quality.

The official API is FP8, which should imply that it's lossless.
Post reply on HN