Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

351–360 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#351

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

What a strangely hostile statement on an open weight model. Running like 20 benchmark evaluations isn't trivial by itself, and even updating visuals and press statements can take a few days at a tech company. It's literally been 5 days since this "new generation" of models released. GPT-5.3(-codex) can't even be called via API, so it's impossible to test for some benchmarks. I notice the people who endlessly praise c…

Isn’t trivial? How is it not completely automated at this point?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#352
post #227

Earlier quoted context omitted.

Thank you for continuing to maintain the only benchmarking system that matters! Context for the unaware: https://simonwillison.net/tags/pelican-riding-a-bicycle/

It's interesting how some features, such as green grass, a blue sky, clouds, and the sun, are ubiquitous among all of these models' responses.

It is odd, yeah.

I'm guessing both humans and LLMs would tend to get the "vibe" from the pelican task, that they're essentially being asked to create something like a child's crayon drawing. And that "vibe" then brings with it associations with all the types of things children might normally include in a drawing.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#353
post #340
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

The bird not having wings, but all of us calling it a 'solid bird' is one of the most telling examples of the AI expectations gap yet. We even see its own reasoning say it needs 'webbed feet' which are nowhere to be found in the image. This pattern of considering 90% accuracy (like the level we've seemingly we've stalled out on for the MMLU and AIME) to be 'solved' is really concerning for me. AGI has to be 100% righ…

MMLU performance caps out around 90% because there are tons of errors in the actual test set. There's a pretty solid post on it here: https://www.reddit.com/r/LocalLLaMA/comments/163x2wc/philip_...

As far as I can tell for AIME, pretty much every frontier model gets 100% https://llm-stats.com/benchmarks/aime-2025

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#354
post #215

Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.

This Pelican benchmark has become irrelevant. SVG is already ubiquitous. We need a new, authentic scenario.

  1. Take the top ten searches on Google Trends 
     (on day of new model release)
  2. Concatenate
  3. SHA-1 hash them
  4. Use this as a seed to perform random noun-verb 
     lookup in an agreed upon large sized dictionary. 
  5. Construct a sentence using an agreed upon stable 
     algorithm that generates reasonably coherent prompts
     from an immensely deep probability space.
That's the prompt. Every existing model is given that prompt and compared side-by-side.

You can generate a few such sentences for more samples.

Alternatively, take the top ten F500 stock performers. Some easy signal that provides enough randomness but is easy to agree upon and doesn't provide enough time to game.

It's also something teams can pre-generate candidate problems for to attempt improvement across the board. But they won't have the exact questions on test day.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#355

Earlier quoted context omitted.

They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.

I think you're seriously underestimating how much effort the fine tuning at their scale takes and what impact it has. They don't pack every edge case into the system prompt either. It's not like they update the model every few hours or even care about memes. If they seriously did, they'd force-delegate spelling questions to tool calls.

Could it be the model is constantly searching its own name for memes, or checking common places like HN and updating accordingly? I have no idea how real-time these things are, just asking.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#356

Earlier quoted context omitted.

Like identifying names of skateboard tricks from the description? https://skatebench.t3.gg/

I couldn't find an about page or similar?

Here's the public sample https://github.com/T3-Content/skatebench/blob/main/bench/tes...

I don't think there's a good description anywhere. https://youtube.com/@t3dotgg talks about it from time to time.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#357
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…

"You're holding it wrong."

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#358

The amount of times benchmarks of competitors said something is close to Claude and it was remotely close in practice in the past year: 0

I honestly feel like people are brainwashed by anthropic propaganda when it comes to claude, I think codex is just way better and kimi 2.5 (and I think glm 5 now) are perfectly fine for a claude replacement.

> I think codex is just way better

Codex was super slow till 5.2 codex. Claude models were noticeably faster.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#359
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

Gemini Pro:

You should definitely drive.

If you walk there, your car will still be dirty back at your house! Since the goal is to get the car washed, you have to take it with you.

PS fantastic question!

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#360
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…

Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.
Post reply on HN