Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

281–290 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#281
post #277

What I haven't seen discussed anywhere so far is how big a lead Anthropic seems to have in intelligence per output token, e.g. if you look at [1]. We already know that intelligence scales with the log of tokens used for reasoning, but Anthropic seems to have much more powerful non-reasoning models than its competitors. I read somewhere that they have a policy of not advancing capabilities too much, so could it be tha…

Intelligence per token doesn't seem quite right to me.

Intelligence per feels closer. Per dollar, or per second, or per watt.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#282

[flagged]

FYI: Chinese models, to be approved by the regulator, have to go through a harness of questions, which of course include this Tiananmen one, and have to answer certain things. I think that on top of that, the live versions have "safeguards" to double check if they comply, thus the freezing.

Unfair competition.

Should western models go through similar regulatory question bank? For example about Epstein, Israel's actions in Gaza, TikTok blocking ICE related content and so on?

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#284
post #238

Earlier quoted context omitted.

I think the only advantage that closed models have are the tools around them (claude code and codex). At this point if forced I could totally live with open models only if needed.

The tooling is totally replicated in open source. OpenCode and Letta are two notable examples, but there are surely more. I'm hacking on one in the evenings. OpenCode in particular has huge community support around it- possibly more than Claude Code.

I know, I use OpenCode daily but it still feels like it's missing something - codex in my opinion is way better at coding but I honestly feel like that's because OpenAI controls both the model and the harness so they're able to fine tune everything to work together much better.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#285
Just tried it, its practically the same as glm-4.7 - it isn't as "wide" as claude or codex so even on a simple prompt is misses out on one important detail - instead of investigating it ploughs ahead with the next best thing it thinks you asked for instead of investigating fully before starting a project.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#286
post #266

Earlier quoted context omitted.

I think the only advantage that closed models have are the tools around them (claude code and codex). At this point if forced I could totally live with open models only if needed.

If tooling really is an advantage why isn't it possible to use the API with a subscription and save money?

In my opinion it is because if you control both the model and the harness then you're able to tune everything to work together much better.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#287
post #233

Earlier quoted context omitted.

They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…

Gemini 3 Pro: This is a classic logistical puzzle! Unless you have a very unique way of carrying your vehicle, you should definitely drive. If you walk there, you'll arrive at the car wash, but your car will still be dirty back at your house. You need to take the car with you to get it washed. Would you like me to check the weather forecast for $mytown to see if it's a good day for a car wash?

For me, various forms of Gemini respond with "Unless you are planning on carrying the car there" which I find to be just sassy enough to be amusing.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#288

What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…

US Secretary of State Bressent just publicly said that the US needs to get along and cooperate with China. His tone was so different than previously in the last year that I listened to the video clip twice.

Obviously for the average US tax payer getting along with China is in our interests - not so much our economic elites.

I use both Chinese and US models, and Mistral in Proton’s private chat. I think it makes sense for us to be flexible and not get locked in.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#289
Why are we not comparing to opus 4.6 and gpt 5.3 codex...

Honestly these companies are so hard to takes seriously with these release details. If it's an open source model and you're only comparing open source - cool.

If you're not top in your segment, maybe show how your token cost and output speed more than make up for that.

Purposely showing prior-gen models in your release comparison immediately discredits you in my eyes.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#290
post #192

Blog post and hugging face link are out. See related thread: https://news.ycombinator.com/item?id=46977210 [1] https://z.ai/blog/glm-5 [2] https://huggingface.co/zai-org/GLM-5

Why did they have to tweak sampling parameters so much for the benchmarks? Looks like rerun hacking.
Post reply on HN