Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

491–500 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#491

Earlier quoted context omitted.

So its like run 3 loops of “here project, find bugs” with all good models, then dedupe and priorize with a sota?

The loop is "look at this file in this repo, find bugs" iterated over every file in a project, with the ability to look at the rest of the repo for cross-file bugs related to the file they're instructed to look specifically at, but yes. The Anthropic folks have basically said that's how they're doing security audits (Nicholas Carlini is an Anthropic employee and he's done talks about it), so I assume that's how Mytho…

Ah this is an important distinction, thanks!

Not sure if helpful but in my experience when something a bit more complex needs to be done, manually making it read the context I know the model will need for it to solve it well (like making it consume all the project docs first) helps with getting a more satisfactory result instead of only giving it the task and let it look around and consume the context it thinks it needs.

Will test your bug finding method in a current project of mine both with my "manual" context preloading and without.

Re: GLM 5.2 beats Claude in our benchmarks

#492

Chinese models are almost certainly cheating on benchmarks, I would bet if you saw the training data that the benchmark canaries are in there. GLM may be a good model in general but it s benchmaxxed and definitely not as good as Opus 4.8.

Why would you say that?

I use DeepSeek V4 Flash (high) and MiMo 2.5 (non Pro, because vision) to work on medium sized projects (~1mil lines of code, C#, Go, TypeScript) with great success.

And that is coming from someone who used Opus 4.7 and GPT 5.5 as workhorses before.

And I'm pretty sure GLM 5.2 is better than the lighter models I use.

My worflow is simple: plan -> clarify -> implement.

1) plan prompt template: I describe what I need and ask LLM to generate a markdown file containing an implementation plan plus at least 10 clarification questions for me to answer.

2) I answer the questions in the plan.md file.

3) implementation prompt template: I ask LLM to implement plan.md and tell me at the end if there were any deviations and new findings during the implementation (there ofter are).

Re: GLM 5.2 beats Claude in our benchmarks

#493
post #262

Earlier quoted context omitted.

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

> but if you only want to use the best model available, it isn't there yet I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way mo…

I would say one thing I've enjoyed about the latest frontier models from US labs is that you just work at a higher level of abstraction. You can talk about the end goal and it'll just rip. You'll add scaffolding to constrain the patterns etc, but I do way less baby sitting than I expected on 5.6 vs 5.4 vs Deepseek v4 Pro.

Re: GLM 5.2 beats Claude in our benchmarks

#494
The title of the post on their blog is really misleading "We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks". Mythos (or Fable) isn't even benchmarked, and there's giant caveat literally at the bottom: "We have a caveat: This is one task, one dataset, one run."

I think the post is still informative, but very a little disingenuous and clickbaity.

Re: GLM 5.2 beats Claude in our benchmarks

#495
post #439

Earlier quoted context omitted.

Except that there is no application for AI in CAD that is better, more appropriate, more robust or more sensible than learning how to use a CAD package and doing it yourself. It's not fast-changing, it's not abstract, it's just not that difficult, and where it is difficult, the AI cannot help you, because it is not capable of things you are capable of. Learn CAD yourself. Honestly; I was sure I would never manage to…

This is a refreshing perspective because recently I feel like I’m surrounded by people who think they can effectively implement complex software, just by hammering the best models. It has been hard to explain that they are in fact just creating toy versions and there is no way they can do it without learning the underlying architecture. But they just keep going wasting 100s of dollars , lost in a sea of bugs

Until a few years ago I'd have been the person who thought you could make a text-to-CAD system scale up to all of it. And then I tried to make stuff I wanted.

Dabbled with OpenSCAD as we will. I decided to learn FreeCAD and what I discovered is that, even putting aside FreeCAD's many documented issues, parametric GUI CAD is not an imprecise, clumsy or fiddly way to work.

It is expressive, precise, generally capable of all the things that code-CAD can do and much more, and it's much, much quicker to work in, once you've learned a few core principles.

As you say, there is an underlying architecture; it's not just a sort of 3D paint package.

The problems the text-as-whatever crowd have are all Dunning-Kruger things in the truest sense.

People who are unaware they are unskilled in a particular technology are unlikely to successfully replace it with another. Particularly one that requires describing the problem domain in precise language.

Quite often when you see text-to-CAD discussions, especially here, there's evidence of profound misunderstandings from the people who think they are going to automate it. They assume their frustrations with the tools stem from limitations of the tools, not from the limits of their understanding.

As a person with decades of experience of code I have found learning how to use LLMs effectively to be much, much harder than learning CAD.

Re: GLM 5.2 beats Claude in our benchmarks

#496
post #355

Earlier quoted context omitted.

I think it would be good not to suggest someone run a new Chinese agent on their bare metal. When I posted the comment I was both the first commentor as well as the first person to upvote the submission. That matters. My name is ALSO on the open source repo that allows Opencode to be run in a container. That's transparency, maybe not here, but on a clickthrough to Github it is immediately obvioius.

> I think it would be good not to suggest someone run a new Chinese agent on their bare metal. Not sure a project nobody knows or uses is much better in this regard?

[flagged]

Re: GLM 5.2 beats Claude in our benchmarks

#497
post #342

Earlier quoted context omitted.

With Claude Code Ultrathink, I used 3 million tokens in 20 minutes. At API prices, that would be around 30$. So 90$/h. Model cost is not that much lower.

x40hrs/week * 50 weeks = $180k Congrats, now you’re paying an engineer’s salary to make your engineer at best 20% more productive. Better to hire another engineer, or two jrs, and build up your in house talent.

Only you get things done lost faster and don't need to pay entire years salary?

Re: GLM 5.2 beats Claude in our benchmarks

#498
post #465

Are open labs just loss leaders backed by Chinese govt? Is this like electric cars where the goal is to flood the market with good enough quality for free so they end up dominating the market? Or is there a business model I’m missing?

US EVs were also heavily subsidized, but they were all built using Chinese parts.

The EV supply chain in the US back in say 2007 certainly had far fewer key parts sourced from China than recent years.

As far as US EVs being subsidized early, if you take state and federal tax incentives, DoE grants and loan guarantees as subsidizes then that's true.

It's debatable (I think incentives applied to all suppliers not just US ones) but a reasonable statement.

Re: GLM 5.2 beats Claude in our benchmarks

#499
post #465

Are open labs just loss leaders backed by Chinese govt? Is this like electric cars where the goal is to flood the market with good enough quality for free so they end up dominating the market? Or is there a business model I’m missing?

US EVs were also heavily subsidized, but they were all built using Chinese parts.

US EVs were "lightly" subsidized compared to what the Chinese govt has done. In the ballpark of 250 billion dollars by the Chinese vs maybe 10% of that by the US.

Re: GLM 5.2 beats Claude in our benchmarks

#500
post #288

Earlier quoted context omitted.

Is there a secure way to use GLM without spending $10K’s for local HW? I “only” have a 128GiB inference machine, and don’t really trust anthropic not to steal my IP over time. I see no reason to trust Z.ai more than other vendors.

You'd need to multiply that $10k by 8 minimum.

[deleted]
Post reply on HN