Live data from Hacker News

GLM 5.2 beats Claude in our benchmarks

semgrep.dev

131–140 of 559 posts

Re: GLM 5.2 beats Claude in our benchmarks

#131
post #24
post #4

Earlier quoted context omitted.

>GLM export controls incoming? US imposing export restrictions on a model from China?

How would that even work for an open-weight model?

Go after the hosts, 99% of people won't be able to run this locally even if they wanted to.

Re: GLM 5.2 beats Claude in our benchmarks

#132
post #6

It reads like an ad. Secondly these are "just" IDORs, arguably the easiest class of vulnerabilities. Thirdly it compares to GPT 5.5 and Opus 4.8. No, we don't have Mythos at home.

In my experience, GLM 5.2 is extremely good at finding vulnerabilities, and more importantly, unlike Opus, I've never seen it refuse a command. It genuinely is a very strong model for finding and fixing vulnerabilities.

More importantly, unlike Mythos and Fable, you can actually use GLM 5.2! It's not just marketingware that got its founder in hot water with the government.

Re: GLM 5.2 beats Claude in our benchmarks

#133

GLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.

Obvious answer: build all your open source LLMs into firearms, get the SC to grant 2A protections.

Re: GLM 5.2 beats Claude in our benchmarks

#134

I like GLM 5.2... ish. It's ok. I'd be mostly fine switching to it. I just can't find a cost effective way to do that. z.AI's coding plan is both overpriced and unreliable. ollama's is also overpriced. Paying by the token for it on openrouter etc is more expensive than just having a Codex or Claude coding plan. If you have to pay by the token, it's clearly cheaper. It's not competitive with a coding plan though.

It also means giving up vision which I don't know how I would deal with. I think I would prefer a weaker model with vision than a stronger without.

Why's that?

Re: GLM 5.2 beats Claude in our benchmarks

#135

Earlier quoted context omitted.

Would you be better off pooling that money with some hackerspace group and then setting up shared inference infra, so that way you at least get better utilization?

And before you know it, you invented some openrouter provider from first principles...

Right. For example you will need to figure out how to share it and who maintains it.

Re: GLM 5.2 beats Claude in our benchmarks

#136

These numbers are seem pretty low compared to what I was able to achieve specifically around windows kernel, win32k win32u to be exact. It honestly wouldn't surprise me anymore if china started surpassing models that US makes public, at least in specific categories such as cyber. GLM 5.2 is already capable enough to assist in self-training which is similar to what we saw happen with frontier models and they appear to…

It will almost for sure surpass the models which Trump will allow US "allies" (which he just considers client states) to use. This, together with China's growing dominance in PV, rechargeable batteries, EV, could really be the nail in the coffin for the post WWII economic world order.

You are delusional if you think China is going to let Europe have access to Mythos level models for free.

Re: GLM 5.2 beats Claude in our benchmarks

#137
post #55

Earlier quoted context omitted.

8 X RTX6000. It will run you around 80-100k to get started with a model at this size with decent tps.. Don't worry though, open source evangelists will tell you that these will be running on your phone in the next 3 years. For $100k you could run this model 24/7 through open router with 10 concurrent sessions at 50tps for a decade and have money left over for a vacation. There's no point in investing this type of mon…

> 8 X RTX6000. It will run you around 80-100k to get started 8 x RTX6000 GPUs cost $100,000 alone. You then need to build a system that can support those GPUs with enough PCIe lanes through a PCIe switch. It's going to be $120K to $150K to build or buy a system to run this.

isn't throwing that into a [insert financial vehicle that gives 99.99999% safe returns] going to destroy that when you factor in electricity costs?

Or even just electricity costs vs token cost

Re: GLM 5.2 beats Claude in our benchmarks

#139
post #59

I have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and Open…

Im really curious about this. Why pay API pricing? I burn 1000s of dollars a month of api according to claude usage but only pay the $100 subscription

Re: GLM 5.2 beats Claude in our benchmarks

#140
post #91
post #58

Earlier quoted context omitted.

you can however, have fun with it. oil workers buy 100k trucks they do not-much with. why not a 100k in computer?

Because car loans can’t be used to buy computers

And there's your idea. If you could find a way to get people to add another $500/month over 80+ months to an auto loan, dealers would eat that up like filet mignon.
Post reply on HN