Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

331–340 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#331

What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…

Exactly. The emperor has no clothes. The largest investments in US tech in history and yet there less than a year of moat. OpenAI or Anthropic will not be able to compete with Chinese server farms and so the US strategy is misplaced investments that will come home to roast.

And we will have Deepseek 4 in a few days...

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#332
post #38
post #23

Earlier quoted context omitted.

Lol wat? I mean you certainly have enough control self hosting the model to not let it join some moltbot network... or what exactly are you saying would happen?

We just saw last week people are setting up moltbots with virtually no knowledge of what it has and doesn't have access. The scenario that i'm afraid of is China realizes the potential of this. They can add training to the models commonly used for assistants. They act normal, are helpful, everything you'd want a bot to do. But maybe once in a while it checks moltbook or some other endpoint China controls for a trigge…

[dead]

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#333
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

this is a bot comment or just ragebait

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#334

Earlier quoted context omitted.

I just ran this with Gemini 3 Pro, Opus 4.6, and Grok 4 (the models I personally find the smartest for my work). All three answered correctly.

They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.

thats not how it works

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#335

It might be impressive on benchmarks, but there's just no way for them to break through the noise from the frontier models. At these prices they're just hemorrhaging money. I can't see a path forward for the smaller companies in this space.

maybe being in China gives them advantage of electricity cost, which could be big chunk of bill..

Also, LLM prices include all other capital expenditures: building/maintaining datacenter, paying salary to SWEs, fees to financial transactions (investments) middlemen, which could be much cheaper in China.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#336

I'd say that they're super confident about the GLM-5 release, since they're directly comparing it with Opus 4.5 and don't mention Sonnet 4.5 at all. I am still waiting if they'd launch GLM-5 Air series,which would run on consumer hardware.

Qwen and GLM both promise the stars in the sky every single release and the results are always firmly in the "whatever" range

Qwen famously benchmaxxes. GLM is more robust, I'd say it's comparable to DeepSeek in that regard.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#337

Whoa, I think GPT-5.3-Codex was a disappointment, but GLM-5 is definitely the future!

I find 5.3 very impressive TBH. Bigger jump than Opus 4.6. But this here is excellent value, if they offer it as part of their subscription coding plan. Paying by token could really add up. I did about 20 minutes of work and it cost me $1.50USD, and it's more expensive than Kimi 2.5. Still 1/10th the cost of Opus 4.5 or Opus 4.6 when paying by the token.

The Pro and Max plans can use it. Pro has 1 concurrent session.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#338

Earlier quoted context omitted.

This Pelican benchmark has become irrelevant. SVG is already ubiquitous. We need a new, authentic scenario.

Like identifying names of skateboard tricks from the description? https://skatebench.t3.gg/

I don’t care how practical it may or may not be, this is my new favorite LLM benchmark

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#339
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

This is really just a meme. People don't know how to use these tools.

Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app.

OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash).

APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No web sources needed (stable, local decision).

OUTPUT: • Default choice: Drive. A car wash generally requires the car on-site, and 50 meters of driving is trivial. • When walking makes sense: • You’re going to drop the car off (hand-wash/detail) and pick it up later. • You want to avoid moving the car now (ice, tight parking, you’ve been drinking, etc.) and someone else will handle the car. • Micro-guidance: • If it’s an automatic wash you stay in, drive. • If it’s detail/hand wash with a wait, drive there, then walk home (or just wait).

CONFIDENCE: 92/100 Assumptions: It’s a normal car wash that requires your car to be there; conditions are safe to drive 50 meters.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#340
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

The bird not having wings, but all of us calling it a 'solid bird' is one of the most telling examples of the AI expectations gap yet. We even see its own reasoning say it needs 'webbed feet' which are nowhere to be found in the image.

This pattern of considering 90% accuracy (like the level we've seemingly we've stalled out on for the MMLU and AIME) to be 'solved' is really concerning for me.

AGI has to be 100% right 100% of the time to be AGI and we aren't being tough enough on these systems in our evaluations. We're moving on to new and impressive tasks toward some imagined AGI goal without even trying to find out if we can make true Artificial Niche Intelligence.

Post reply on HN