Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

411–420 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#411

Earlier quoted context omitted.

This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…

Interesting, what were the instructions if you don't mind sharing?

[deleted]

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#412
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

What a weird thing to say considering humans have tons of blind spots and missing knowledge, do dumb things, make easy to miss mistakes. I guess they lack intelligence too.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#413
post #69
post #5

It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…

I don't really care about being able to self host these models, but getting to a point where the hosting is commoditised so I know I can switch providers on a whim matters a great deal. Of course, it's nice if I can run it myself as a last resort too.

It is pretty easy to set up Open Router and set up schemes to point at different models, but in the same token, you can point at yours locally unless you wanted a "more powerful" answer

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#414
post #373

Earlier quoted context omitted.

Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.

That's the answer of his LLM which has decomposed the question and built the answer following the op prompt obviously. I think you didn't get it.

> I think you didn't get it.

I did get it, and in my view my point still stands. If I need to use special prompts to ask such a simple question, then what are we doing here? The LLMs should be able to figure out a simple contradiction in the question the same way we (humans) do.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#415

So that was pony alpha (1). Now what's Aurora Alpha? (1) https://openrouter.ai/openrouter/pony-alpha

It's GPT. Tried and reproduced some polluted single-token Chinese phrases from 4o era.

It certainly likes producing long responses littered with markdown tables like GPT. Not quite as verbose as the gpt-5 family, though.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#416

Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.

This aligns very closely with my experience. When left to its own devices, GLM-4.7 frequently tries to build the world. It's also less capable at figuring out stumbling blocks on its own without spiralling. For small, well-defined tasks, it's broadly comparable to Sonnet. Given how incredibly cheap it is, it's useful even as a secondary model.

How is the web search functionality? I have only used deepseek to lower costs from gpt api but had to incorporate a serper to actually do web searches

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#417
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

In some ways, Opus 4.6 is a step backwards due to massively higher token consumption.

yeah, I am still using 4.5 for coding.

I have started using Gemini Flash on high for general cli questions as I can't tell the difference for those "what's the command again" type questions and it's cheap/fast/accurate.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#418
post #233

The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

They all get it right if you allow them to think.

I just copy pasted your question "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" without any further prompt and ran it against GLM 5, GPT 5.2, Opus 4.6, Gemini 3 Pro Preview, through OpenRouter with reasoning effort set to xhigh.

Not a single one said I should walk, they all said to drive.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#419
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…

None of that stuff is necessary, they all get it right with the initial question and no further prompt if you dial the reasoning effort up.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#420
post #410

Earlier quoted context omitted.

If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic. What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how . I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and…

I'm not looking for full automation But your boss probably is.

Full automation is also possible by putting your coding agent into a loop. The point is that an LLM that can solve a small task is more valuable for quality output, than an LLM that can solve a larger task autonomously.
Post reply on HN