Live data from Hacker News

GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

z.ai

451–460 of 540 posts

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#451

Earlier quoted context omitted.

I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…

Can you use the models you get through Gemini Ultra in Claude Code? If not, what coding tool do you use?

Claude code router

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#452
post #365
post #326

Earlier quoted context omitted.

They will start to max this benchmark as well at some point.

It's not a benchmark though, right? Because there's no control group or reference. It's just an experiment on how different models interpret a vague prompt. "Generate an SVG of a pelican riding a bicycle" is loaded with ambiguity. It's practically designed to generate 'interesting' results because the prompt is not specific. It also happens to be an example of the least practical way to engage with an LLM. It's no mo…

RLHF (reinforcement learning from human feedback) is to a large extent about resolving that ambiguity by simply polling people for their subjective judgement.

I've worked one an RLHF project for one of the larger model providers, and the instructions provided to the reviewers were very clear that if there was no objective correct answer, they were still required to choose the best answer, and while there were of course disagreements in the margins, groups of people do tend to converge on the big lines.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#453

GLM-4.7-Flash was the first local coding model that I felt was intelligent enough to be useful. It feels something like Claude 4.5 Haiku at a parameter size where other coding models are still getting into loops and making bewilderingly stupid tool calls. It also has very clear reasoning traces that feel like Claude, which does result in the ability to inspect its reasoning to figure out why it made certain decisions…

I'm not sure what it is about GLM 4.7 Flash, but it definitely seems to nail a sweet spot. Even the supposedly frontier models make a mess of large requests, so small, well-scoped requests are the way, IMO; and in that space, 4.7 Flash holds its own better than it has any right to.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#454
post #381

I've been using GLM 4.7 with opencode. It is for sure not as good but the generous limits mean that for a price I can afford I can use it all day and that is game changer for me. I can't use this model yet as they are slowly rolling it out but I'm excited to try it.

Exactly! I don't understand comments claiming GLM-4.7 is very bad.

With New Year's promotional discount, I got Lite coding version for ~3$ per month. I have burned couple dozen million of tokens in a session and 5h allowance barely budged. For what I do on personal time - I will never burn through it[0].

I have Claude Code Opus 4.6 at work - yes GLM-4.7 is not as good, though for personal work on bootstraping some applications - it's excellent.

I feel like it's literally 6-9 months behind SOTA, most expensive LLM tools that my employer was buying for me and my colleagues, for 3$ per month (even if it's 10$ without discount). Will see how it's with GLM-5 when Z.AI lite coding plan will get it, but I feel the gap to SOTA is narrowing and fast.

[0] Though I feel like a stone age neanderthal, when people say they run multiple agents in parallel and burn tens of millions of tokens in minutes.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#455
I paid for the $30 plan. It's useful to me via OpenCode as a cheap backend for CLI/Agentic workflows.

I also want to try it with Wiggam Loop to test whether they can together build production-level code if guided via prompts and a PRD. Let's see!

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#456

Earlier quoted context omitted.

Just to say - 4.6 really shines on working longer without input. It feels to me like it gets twice as far. I would not want to go back.

If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic. What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how . I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and…

That sounds like wishful thinking. Every client I work for wants to reduce the rate at which humans need to intervene. You might not want that, but odds are your CEO does. And babysitting intermediate stages is not productive use of developer time.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#457

Earlier quoted context omitted.

Their 50 USD per month plan gives you 24M tokens per day: https://www.cerebras.ai/pricing

I had that for a few months and cancelled. They have minutely rate limits as well so you get 3-4 hyperspeed responses and then a 45 second pause waiting for the throttling to let your next request through. And then, depending on what you're working on, the 24M daily allotment is gone in under an hour. I regularly burned it in about 25 minutes of agent use. I imagine if I had infinite budget to pay regular API rates o…

> They have minutely rate limits as well so you get 3-4 hyperspeed responses and then a 45 second pause waiting for the throttling to let your next request through.

I haven’t really gotten that, though have noticed on some occasions:

A) high server load notifications, most commonly, can delay an answer by about 3-10 seconds

B) hangs, this happens quite rarely, not sure if a network issue or something on their side, but sometimes the submitted message just freezes (e.g. nothing happening in OpenCode), doesn’t seem deliberate because resubmitting immediately works, more often than not

> And then, depending on what you're working on, the 24M daily allotment is gone in under an hour. I regularly burned it in about 25 minutes of agent use.

That’s a lot of tokens, almost a million a minute! Since the context is about 128k, you’d be doing about 8 full context requests every minute for 25 minutes straight.

I can see something like that, but at that point it feels like the only thing that’d actually be helpful would be caching support on their end.

You must be on some pretty high tier subscriptions with the other providers to get the same performance!

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#458
post #29

Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…

"the greatest theft in human history" what a nonsense. I was curious, how the AI haters will cope, now that the tides here have changed. We have built systems that can look at any output and replicate it. That is progress. If you think some particular sequence of numbers belongs to you, you are wrong. Current intellectual property laws are crooked. You are stuck in a crooked system.

Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks

#460

Earlier quoted context omitted.

Anthropic has very tight limits, so you're basically using the worst (pricing-wise) SOTA cloud model as your baseline. I have $200 subs for both Claude and OpenAI, and I also bump into limits with Claude all the time, whether coding or research. With Codex, I ran into the limit once so far, and that's in a month of very heavy (sometimes literally 24 hours around the clock, leaving long-running tasks overnight) use.

I bought the Gemini Ultra to try for a month (at the discounted price). I have been using it non-stop for Opus 4.6 Thinking, which is much better than Gemini 3 Pro (High) and it's been a blast. The most I've managed to consume is 60% of my 5 hourly quota. That was with 2-3 instances in parallel. I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and go…

How do you use Opus through Gemini Ultra? I must be missing something
Post reply on HN