Live data from Hacker News

GLM-5.2 is a step change for open agents

interconnects.ai

71–80 of 240 posts

Re: GLM-5.2 is a step change for open agents

#71
post #44

Earlier quoted context omitted.

With subsidization from the Chinese government they will probably be equal to or better than the models here. I mean, have you looked at the author list of any given AI paper published within, say, the past 5 years? I wouldn't be surprised if half or more AI researches are from China.

Can you compare the amount to the USA subsidization? Which one is bigger? Per Capita? Per unit of economic growth achieved?

You mean from the private investors? It seems the labs on both sides of the ocean are quite negative in their profitability right now due to the competitiveness. Though Anthropic claims they will have a profitable quarter this year (despite the huge build-out), so their margins on API costs are likely quite decent.

Re: GLM-5.2 is a step change for open agents

#72

Earlier quoted context omitted.

Yeah. There's no way to verify what these providers are doing. The real future is running these models at home. Opus level inference on our own hardware would be a dream come true.

How will anyone running home instances be able to compete against people paying some money running much more powerful models on much more powerful hardware?

It’ll be interesting.

I’m using Qwen3.6:27B at home and mostly Sonnet/Opus (depending on the complexity of the task) at work.

You have to break things down into smaller chunks for the local models. For the bigger cloud ones they can do a lot of the broader thinking.

Re: GLM-5.2 is a step change for open agents

#73
post #50

GLM-5.2 has been a step change in how fast i can burn through tokens. I subscribed to their max plan to try it out. It counted me 700M tokens and drained my weekly quota in under 2 days. Quota just reset less than 24h ago and i'm already >60% weekly quota usage. For reference the kind of work i did would have used somewhere between 3% and 5% of Codex max or Claude max. The model is good, the plan is a scam

What kind of tasks have you been using it for?

Re: GLM-5.2 is a step change for open agents

#75

Here are the numbers from their bar chart: 1. SWE-bench Pro Model Score (%) GLM-5.2 62.1 GLM-5.1 58.4 Claude Opus 4.8 69.2 GPT-5.5 58.6 Gemini 3.1 Pro 54.2 2. Terminal-Bench 2.1 Model Score (%) GLM-5.2 81.0 GLM-5.1 63.5 Claude Opus 4.8 85.0 GPT-5.5 84.0 Gemini 3.1 Pro 74.0 3. NL2Repo Model Score (%) GLM-5.2 48.9 GLM-5.1 42.7 Claude Opus 4.8 69.7 GPT-5.5 50.7 Gemini 3.1 Pro 33.4 4. DeepSWE Model Score (%) GLM-5.2 46.2…

copying the graphs and tables to HN is noisy and harder to read

Re: GLM-5.2 is a step change for open agents

#76

I can't help wondering what kind of models we'll see coming out of China once it gets its own chip fabs up and running. Right now it sounds like the US's export ban is not slowing them down a whole lot.

> Right now it sounds like the US's export ban is not slowing them down a whole lot. It may wind up being a massive boost to them in the long run, even. Necessity is the mother of invention.

Trump allowed more advanced chips (H200s) to be sold after his visit, because some people in the admin still believe the US can "addict" China to the hardware. It seems China is only letting a token few in, the ban is more on their side now, as Xi really wants indiginous capability.

Re: GLM-5.2 is a step change for open agents

#78

I've been working with Deepseek V4 Flash (with opencode as the harness). It's been almost indistinguishable from Codex / Claude Code for me. I'm sure I'll run into problems when I get to a stickier ticket to tackle. But so far, it's been quite good, and I find it writes straightforward code. I do think the Chinese models are good enough for an 80/20 rule use case.

I tried Deepseek V4 Flash with very low expectations and was pleasantly surprised. It's a surprisingly capable model for the price.

Re: GLM-5.2 is a step change for open agents

#79
post #19

Earlier quoted context omitted.

$200/h is on the extreme end and I would argue most people here aren't anywhere close to that. The median hourly wage in the US is $28/h, this equates to nearly 7.5 hours. A full day of work a month for the average person to use Claude with reasonable limits. Yes, the people on $28/h may not be the software development types, so their income might not be as high, but these are the people who would probably be vibe co…

I suspect the reply above is referring to charge out rates rather than wages.

My fault, thanks for the correction (:

Re: GLM-5.2 is a step change for open agents

#80
post #69

Earlier quoted context omitted.

If you are running multiple agents your cost to them should be multiples less what their roi is.

My costs are 0$ as any token or subscription spend on agents is invoiced as an expense to my clients.

Thanks so much for being bold enough to be fairly open about the costs, how you arrange billing and the advantages that's given you.

I've been fooling around with DeepSeek 4 agentically. It's probably not as good as Anthropic offerings, but even those seem to be roiled in politics and strife and DeepSeek 4 is very good IMHO. I'll later try out GLM.

I'm in Australia. The government has set up a "return and earn" scheme to keep aluminium cans, plastic bottles and paper drink cartons out of the waste stream. A laudable project. The money you make from return drink containers is pretty low, $AU 0.1 per container. I've participated to get the rubbish out of natural water streams and to make a nano amount of money on the side.

When I looked at the costs of an app I was getting DeepSeek to help me with, I realised that the several hours I'd spent learning and building had cost something like 8 recycled containers. In my head after doing some DeepSeek stuff, I calculate a "cans per app" metric for myself for fun. I may even setup a simple graph to view my costs that way.

I kind of hope the Anthropics of the world get enough price competition from sources like DeepSeek and GLM to drop their prices significantly. Time will tell.

I'm using the Chinese DeepSeek provider, so everything done there could potentially be taken and used by the CCP... But this is hobbyist learning.

There is probably a market for Deepseek/GLM served from non CCP available servers. I might even look into how hard that would be to setup here.

I also hope that inference focused hardware will come to the fore, reducing energy use and cost. Realistically this will take time though, on the order of years.

Here in Oz, we have community batteries that community members can charge and later draw from. Their electricity prices are competitive. I wonder if someone could setup something like a community battery to run data centres... That way reasonable environmental consideration could be given to inference power generation... This might not work in a market like the US or Europe, but small market size might be an advantage... Who knows.

Post reply on HN