I got fed up with GLM-4.7 after using it for a few weeks; it was slow through z.ai and not as good as the benchmarks lead me to believe (esp. with regards to instruction following) but I'm willing to give it another try.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
191–200 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#192See related thread: https://news.ycombinator.com/item?id=46977210
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#193Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#194Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#195Here is the pricing per M tokens. https://docs.z.ai/guides/overview/pricing Why is GLM 5 more expensive than GLM 4.7 even when using sparse attention? There is also a GLM 5-code model.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#196Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#197Whoa, I think GPT-5.3-Codex was a disappointment, but GLM-5 is definitely the future!
But this here is excellent value, if they offer it as part of their subscription coding plan. Paying by token could really add up. I did about 20 minutes of work and it cost me $1.50USD, and it's more expensive than Kimi 2.5.
Still 1/10th the cost of Opus 4.5 or Opus 4.6 when paying by the token.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#198Whoa, I think GPT-5.3-Codex was a disappointment, but GLM-5 is definitely the future!
Care to elaborate more?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#199The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
Particularly for tool use.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#200Earlier quoted context omitted.
You're surprised that chinese model makers try to follow chinese law?
This is a classic test to see if the model is censored, as censorship is rarely limited to just one event, which begs the question: what else is censored or outright changed intentionally?
You tell me which one is less censored & more trustworthy from those 20,000 killed children's point of view.