Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

481–490 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#481
post #439

Earlier quoted context omitted.

I agree, but there are use cases for the 'best model' other than converting your 1975 stuff to rust: for use cases where LLMs are just getting started to be useful I really want to use the current 'best' model: e.g. CAD, PCB design etc. In particular anything which requires spatial reasoning. The short time I had access to Fable 5 - it was just way better than any other model.

Except that there is no application for AI in CAD that is better, more appropriate, more robust or more sensible than learning how to use a CAD package and doing it yourself. It's not fast-changing, it's not abstract, it's just not that difficult, and where it is difficult, the AI cannot help you, because it is not capable of things you are capable of. Learn CAD yourself. Honestly; I was sure I would never manage to…

This is a refreshing perspective because recently I feel like I’m surrounded by people who think they can effectively implement complex software, just by hammering the best models.

It has been hard to explain that they are in fact just creating toy versions and there is no way they can do it without learning the underlying architecture. But they just keep going wasting 100s of dollars , lost in a sea of bugs

Re: GLM 5.2 beats Claude in our benchmarks

#482
post #170
post #59

I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and Open…

Twenty dollars? How are you comfortable spending that much to write something as simple as a matrix bot? Are people doing this kind of thing just super rich or am I missing something?

Many factor to consider, really, but if it can build be a project while I'm in gym or walking around the city with my Fujifilm - 20$ is a good trade.

Re: GLM 5.2 beats Claude in our benchmarks

#483
post #127

Earlier quoted context omitted.

Quantizing is one thing. But in general it's self-evident that training the model on information that is irrelevant to your use case does not necessarily improve ability, otherwise you'd have AGI just from reinforcing your model on memorizing the first 10^50 digits of pi. Likewise, LLMs do not violate the laws of information theory, and therefore the only way to encode X amount of information in Y amount of bits wher…

> it's self-evident that training the model on information that is irrelevant to your use case does not necessarily improve ability We don’t understand AI or natural intelligence well enough to make such statements. As for self evidence, cross-domain competence in humans and the rise of generalist models over domain-specific ones (on competence, not cost) seems to pretty directly tank your hypothesis.

> We don’t understand AI or natural intelligence well enough to make such statements.

If you believe this then you don't understand AI or natural intelligence well enough to refute my statements either.

Perhaps you're trying to refer to something specific by "cross-domain" competence, but firstly, humans vastly overestimate the extent to which experts in one domain can be trusted to speak accurately on topics in other domains (this is a form of authority bias), and secondly, real cross-domain expertise is a result of pre-existing metacognitive ability such as keen reasoning ability, intense focus, and learning-how-to-learn. In other words, Leonardo da Vinci was not a genius because he was a polymath; he was a polymath because he was a genius.

Likewise, I see no evidence that "generalist models" have proven anything about their ability over domain-specific ones other than that the big AI firms seem to believe that "generalist models" are their golden ticket to AGI and therefore a quintillion-dollar valuation. It's obvious in the long run that tools built for specialized tasks will outperform generalist tools for specific tasks, in the same way that a multi-axis CNC mill does not outperform your bog-standard lathe for shaping objects with rotational symmetry, or perhaps more pertinently to this conversation, how no LLM will ever outperform Stockfish at chess.

Re: GLM 5.2 beats Claude in our benchmarks

#485
post #446
post #143

Earlier quoted context omitted.

Anyone done any benchmarks on the NV4FP quant? Seriously considering pitching an 8 x RTX 6000 Pro box at work to run GLM-5.2 in an air gapped environment.

At that price point you could also go with a Tenstorrent Galaxy Blackhole, which starts at $110,000.

Ooh, I hadn't seen these yet! That looks quite compelling, my only hesitancy would be what the software support looks like. But 1 TB of memory for $110k is really intriguing - I might go bother a sales rep. Thanks!

Re: GLM 5.2 beats Claude in our benchmarks

#487

Are open labs just loss leaders backed by Chinese govt? Is this like electric cars where the goal is to flood the market with good enough quality for free so they end up dominating the market? Or is there a business model I’m missing?

> Are open labs just loss leaders backed by Chinese govt

There are many layers of Chinese govt. But GLM is backed by Beijing municipal govt and Tsinghua University.

Re: GLM 5.2 beats Claude in our benchmarks

#488
If it’s not quite as good as the hype yet, I expect it probably will be in the near future. To do a lot of the primary coating tasks needed for most situations, it’s probably gonna be good enough if it isn’t ready. The harness will be there as well.

Re: GLM 5.2 beats Claude in our benchmarks

#489
post #263

Earlier quoted context omitted.

Good luck. I’m in the legal field, and even there, selling airgapped is tough.

What are the challenges you've seen in selling air gapped? Is it the high upfront cost? Challenges with hardware maintenance or something else?

We already use AWS. Everyone else is using AWS, so if there's an issue we can just say we were following industry standards.

Re: GLM 5.2 beats Claude in our benchmarks

#490

I am using this with a workflow of Claude Code, Codex, Kimi and GLM and the results are pretty astounding and almost 90% of the times Claude's findings and plans are overturned with Claude's agreement.

Exactly the same i am now trying to use and will keep you updated
Post reply on HN