Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

321–330 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#321
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

How many pelican riding bicycle SVGs were there before this test existed? What if the training data is being polluted with all these wonky results...

I'd argue that a models ability to ignore/manage/sift through the noise added to the training set from other LLMs increases in importance and value as time goes on.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#322
post #205

Earlier quoted context omitted.

It's roughly three times cheaper than GPT-5.2-codex, which in turn reflects the difference in energy cost between US and China.

It reflects the Nvidia tax overhead too.

Not really, Western AI companies can set their margins at whatever they want.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#323
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

This Pelican benchmark has become irrelevant. SVG is already ubiquitous. We need a new, authentic scenario.

Like identifying names of skateboard tricks from the description? https://skatebench.t3.gg/

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#324

What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…

US Secretary of State Bressent just publicly said that the US needs to get along and cooperate with China. His tone was so different than previously in the last year that I listened to the video clip twice. Obviously for the average US tax payer getting along with China is in our interests - not so much our economic elites. I use both Chinese and US models, and Mistral in Proton’s private chat. I think it makes sense…

>His tone was so different than previously in the last year that I listened to the video clip twice.

US bluff got called. A year back it looked like US held all the cards and could squeeze others without negative consequences. i.e. have cake and eat it too

Since then: China has not backed down, Europe is talking de-dollarization, BRICS is starting to find a new gear on separate financial system, merciless mocking across the board, zero progress on ukraine, fed wobbled, focus on gold as alternate to US fiat, nato wobbled, endless scandals, reputation for TACO, weak employment, tariff chaos, calls for withdrawal of gold from US's safekeeping, chatter about dumping US bonds, multiple major countries being quite explicit about telling trump to get fucked

Not at all surprised there is a more modest tone...none of this is going the "without negative consequences" way

>Mistral in Proton’s private chat

TIL

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#325

Earlier quoted context omitted.

Sure. My sole point is that calling Opus 4.5 and GPT-5.2 "last generation models" is discounting how good they are. In fact, in my experience, Opus 4.6 isn't much of an improvement over 4.5 for agentic coding. I'm not immediately discounting Z.ai's claims because they showed with GLM-4.7 that they can do quite a lot with very little. And Kimi K2.5 is genuinely a great model, so it's possible for Chinese open-weight m…

I think there are two types of people in these conversations: Those of us who just want to get work done don't care about comparisons to old models, we just want to know what's good right now. Issuing a press release comparing to old models when they had enough time to re-run the benchmarks and update the imagery is a calculated move where they hope readers won't notice. There's another type of discussion where some…

It's high-interest to me because open models are the ultimate backstop. If the SOTA hosted models all suddenly blow up or ban me, open models mitigate the consequence from "catastrophe" to "no more than six to nine months of regression". The idea that I could run a ~GPT-5-class model on my own hardware (given sufficient capex) or cloud hardware under my control is awesome.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#326
post #227
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

Thank you for continuing to maintain the only benchmarking system that matters! Context for the unaware: https://simonwillison.net/tags/pelican-riding-a-bicycle/

They will start to max this benchmark as well at some point.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#327
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

It's unclear where the car is currently from your phrasing. If you add that the car is in your garage, it says you'll need to drive to get the car into the wash.

Do you think the average person would need this sort of clarification? How many of us would have recommended to walk?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#328

Earlier quoted context omitted.

I just ran this with Gemini 3 Pro, Opus 4.6, and Grok 4 (the models I personally find the smartest for my work). All three answered correctly.

They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.

I think you're seriously underestimating how much effort the fine tuning at their scale takes and what impact it has. They don't pack every edge case into the system prompt either. It's not like they update the model every few hours or even care about memes. If they seriously did, they'd force-delegate spelling questions to tool calls.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#329
Been playing with it in opencode for a bit and pretty impressed so far. Certainly more of an incremental improvement than a big bang change, but it does seem better a good bit better than 4.7, which in turn was a modest but real improvement over 4.6.

Certainly seems to remember things better and is more stable on long running tasks.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#330

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

Anthropic, OpenAI and Google have real user data that they can use to influence their models. Chinese labs have benchmarks. Once you realize this, it's obvious why this is the case. You can have self-hosted models. You can have models that improve based on your needs. You can't have both.

zAI, minimax and Kimi have plenty of subscriber usage on their own platforms. They get real data just as well. Less or it maybe but it's there.
Post reply on HN