Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

201–210 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#201

Earlier quoted context omitted.

The Americans may ban the use of the Chinese models in America. But like the Chinese car ban, everyone else will use them.

That's not necessarily a good thing for everyone else, mind. Yes, you get your free model, but the cost of this is not developing your own capability and tying your fate to a country which may or may not have your best interests as a nation in mind. This is just the deindustrialization that occurred in my home region (the American Midwest) playing out on a global scale in different sectors. It was originally driven b…

So... how's that any different from using American stuff for those of us in the rest of the world?

Over the last decade, the US has been way more unreliable than China. There's been a near constant negative impact from the US doing something.

At least with China, we are very good at winning trade wars with them here in Australia.

Re: GLM 5.2 beats Claude in our benchmarks

#202
post #190

Earlier quoted context omitted.

I ran it on my laptop, which is a Lenovo Legion 5i (think 32 GB RAM, 4060 w/ 8 GB VRAM, you get the picture). It was a quantized model (otherwise it would not fit on my NVMe 1TB drive) at 4 bits per weight - UD_Q4_K_XL. It ran at about 12 seconds per token (not tokens per second). A fun project, but not worth it. I used 4096 tokens of context cache, and I ran it with llama.cpp - as it supports memory mapping. Because…

Thank you for sharing. 1.5TB of streamed data at 12 seconds per token on a high end consumer laptop is a pretty high requirement - I can only imagine how much that cost to train. I don't know how running this model could be cost effective for anybody.

Indeed - definitely not cost effective to run it on this laptop LOL. It makes me wonder how fast we could run the model if we could fit the weights entirely within CPU cache (assuming a whole ton of CPUs with low latency & high speed IO of course).

Re: GLM 5.2 beats Claude in our benchmarks

#204
post #193

Earlier quoted context omitted.

It'd be less about "safety" and more "we've spent trillions developing these AI tools only to have the Chinese, once again, copy them and offer them for pennies on the dollar, and no one seems to care about the impact that has on the long-term sustainability of this sector of the American economy as a whole, so we're yanking the models."

"I'm going to take this box razor and make some really deep cuts around the middle of my face because my tech sector is too good and that's actually a bad thing because $foreigners ."

I'm not saying it's necessarily a good thing. I'm also not saying it's about foreigners at this point. It's about seeing a bet through. They've burned a metric crapload of capital on developing AI models and the infrastructure to host them. They want that money back and then some. Remember, the fine shareholders of OpenAI think that 100x returns just aren't reasonable and want that cap lifted. If this kind of thing continues, they'd be lucky to make their money back at all, let alone 100x.

Which would be fine, but as we know, people securitize the crap out of their investments these days, and least some people probably leveraged themselves on some US AI companies, so now the risk is spreading outside of the sector to the economy in general, which is made worse by the sheer amount of spending on AI.

Re: GLM 5.2 beats Claude in our benchmarks

#205
post #59

I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and Open…

GLM 5.2 is a great model, but if you only want to use the best model available, it isn't there yet. Every lab releases models that memorize benchmark answers, both intentionally and unintentionally. But we consistently find that models from Chinese labs have a wider gap between public benchmarks and our evaluations, which we designed to be less vulnerable to benchmaxxing. In multi-agent coding environments, GLM 5.2 i…

Why Deepseek v4 flash is better than pro in your benchmarks?

Re: GLM 5.2 beats Claude in our benchmarks

#206
post #8

Earlier quoted context omitted.

While unlikely , it is not without precedent , there are restrictions on ASML a Dutch company to sell EUV machines

ASML complies as an ally, why would China comply? The weights are already available and downloaded, is it going to be a crime to have them, run them, make them available? Constitutional rights still exist (I hope)

That too has precedence , there is long history of controls of cryptographic algorithms up until the 90s. It wasn't abstract either, older greybeards would remember browsers like Netscape had two versions International and U.S. for this reason.

If you classify AI as a weapon which seems to be the direction that we are all heading towards, they yes first amendment rights won't likely apply.

Re: GLM 5.2 beats Claude in our benchmarks

#207
post #59

I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and Open…

Im really curious about this. Why pay API pricing? I burn 1000s of dollars a month of api according to claude usage but only pay the $100 subscription

[dead]

Re: GLM 5.2 beats Claude in our benchmarks

#208
post #127

Earlier quoted context omitted.

Quantizing is one thing. But in general it's self-evident that training the model on information that is irrelevant to your use case does not necessarily improve ability, otherwise you'd have AGI just from reinforcing your model on memorizing the first 10^50 digits of pi. Likewise, LLMs do not violate the laws of information theory, and therefore the only way to encode X amount of information in Y amount of bits wher…

> But in general it's self-evident that training the model on information that is irrelevant to your use case does not necessarily improve ability, otherwise you'd have AGI just from reinforcing your model on memorizing the first 10^50 digits of pi. It's hardly self-evident, and your counter-example is hardly applicable. The first 10^50 of pi is not the same as having BREADTH of information in the training data, whic…

It is self-evident. Bringing up Kolmogorov complexity is irrelevant, we're talking about rote memorization, but if you can't ignore the given example then replace "digits of pi" with "bits of output from a true random number generator". There's an infinite amount of information that we could shove into a model, and a finite amount of bits with which to store any of that information such that it can be usefully recalled or form useful logical associations.

Re: GLM 5.2 beats Claude in our benchmarks

#209

Earlier quoted context omitted.

That's not necessarily a good thing for everyone else, mind. Yes, you get your free model, but the cost of this is not developing your own capability and tying your fate to a country which may or may not have your best interests as a nation in mind. This is just the deindustrialization that occurred in my home region (the American Midwest) playing out on a global scale in different sectors. It was originally driven b…

So... how's that any different from using American stuff for those of us in the rest of the world? Over the last decade, the US has been way more unreliable than China. There's been a near constant negative impact from the US doing something. At least with China, we are very good at winning trade wars with them here in Australia.

You might feel differently if you were a Filipino or Vietnamese fisherman whose family relied on the income from the stocks of the South China Sea, or a Uighur person living in Western China, or a Ukrainian soldier who has to deal with drones built with Chinese components, or a democracy advocate in Hong Kong, or arguably, a person who had plans for 2020-2021.

Or, on a more local note, an Australian automotive worker who worked for a company that figured out 10 years ago that they wouldn't be able to pay him a decent wage, compete with the then-upcoming Chinese EVs, and remain profitable.

Re: GLM 5.2 beats Claude in our benchmarks

#210
Many people here are now realizing that open weight models are now able to compete against frontier closed models.

This is where we are heading and why many closed labs are terrified of this affecting their bottom line and the reason why they want them banned from being released.

Post reply on HN