Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

261–270 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#261

Earlier quoted context omitted.

Yeah it seems they did not align it to much, at least for now. Yesterday it helped me bypass the bot detection on a local marketplace. that i wanted to scrap some listing for my personal alerting system. Al the others failed but glm5.1 found a set of parameters and tweaks how to make my browser in container not be detected.

Model doing what the user wants with high quality is definitely aligned in my book.

This can never go wrong!

Re: GLM-5.1: Towards Long-Horizon Tasks

#262
post #164

Earlier quoted context omitted.

Computers get better and cheaper. That’s not a forever problem.

Source? GPU and RAM prices have definitely not made consumer PC's cheaper than they were before bitcoin blew up or before AI blew up. Maybe you could make an argument that they are more cost efficient for the price point... But that's not the same as cheaper when every application or program is poorly optimized. For example why would a browser take up more than a GB or two of RAM? And I'd postulate that R&D to develo…

Moore's Law.

We've had RAM shocks before. We nerds can't control the Wall Street or Virginians who like to break the world every so often for the lulz. However, a wobble on the curve doesn't change the curve's destination.

Re: GLM-5.1: Towards Long-Horizon Tasks

#263
post #164

Earlier quoted context omitted.

Computers get better and cheaper. That’s not a forever problem.

Source? GPU and RAM prices have definitely not made consumer PC's cheaper than they were before bitcoin blew up or before AI blew up. Maybe you could make an argument that they are more cost efficient for the price point... But that's not the same as cheaper when every application or program is poorly optimized. For example why would a browser take up more than a GB or two of RAM? And I'd postulate that R&D to develo…

You have to look a bit more long term? 256Mb of what today is slow af RAM used to be pretty pricey. Price will pullback.

Re: GLM-5.1: Towards Long-Horizon Tasks

#264
post #176
post #154

Earlier quoted context omitted.

The 4-bit quants are 350GB, what hardware are you talking about?

qwen3:0.6b is 523mb, what model are you talking about? You seem to have a specific one in mind but the parent comment doesn't mention any. For a hobby/enthusiast product, and even for some useful local tasks, MoE models run fine on gaming PCs or even older midrange PCs. For dedicated AI hardware I was thinking of Strix Halo - with 128gb is currently $2-3k. None of this will replace a Claude subscription.

> qwen3:0.6b is 523mb, what model are you talking about?

1) What are you going to use that for? 0.6 model gives you what you could get from Siri when it first launched at most unless you do some tunning.

2) Pretty clear that they are talking about GLM-5.1 4-bit quant.

Re: GLM-5.1: Towards Long-Horizon Tasks

#265
post #247

Earlier quoted context omitted.

“GLM5…better than Opus, Codex, Gemini…” What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided. Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?

I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…

It’s not a flame war, and you’re not just sharing your experience and encouraging others to try it out.

You’re making a claim, and I’m pointing out that it’s unsubstantiated and not consistent with any other source of data, including that internal to the company that makes the model.

I hope you can see that that’s different than saying it’s worked well for me

Re: GLM-5.1: Towards Long-Horizon Tasks

#266
post #235
post #221

Earlier quoted context omitted.

If it's relevant to the discussion, I hope not. I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation. Let me know if you know of a better platform you can use…

It's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...

This one's even more interesting

https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu...

Who knew Anthropic was this far behind???

Re: GLM-5.1: Towards Long-Horizon Tasks

#267
post #247

Earlier quoted context omitted.

I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…

It’s not a flame war, and you’re not just sharing your experience and encouraging others to try it out. You’re making a claim, and I’m pointing out that it’s unsubstantiated and not consistent with any other source of data, including that internal to the company that makes the model. I hope you can see that that’s different than saying it’s worked well for me

Sometimes we STEM folks are way too rigid, I obviously meant "IN MY OPINION, GLM models are at this point superior to...".

I do not think that anyone who read my comment understood it differently. But I grant you this point, this is just my opinion based on my personal experience not the result of a scientific study.

Once this is said, i wasn't submitting a scientific paper for preprint, just posting my opinion on an internet forum.

Not sure why you are making such a big deal out of it, especially for something for which people can decide within minutes if it works for them or not. And I haven't seen you nitpick on other people saying that all Chinese models are garbage incapable of doing even the most basic task, without quoting any study. This kind of scrutiny tends to be one-sided.

Edit: and regarding what the z.ai team is saying about their models, just check their Discord and the articles they link there. They themselves say that their latest models have leading performance on a number of aspects. It is misleading to suggest that the authors of the model are not proudly saying that their models have best in class performance.

Re: GLM-5.1: Towards Long-Horizon Tasks

#268
post #235

Earlier quoted context omitted.

It's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...

This one's even more interesting https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu... Who knew Anthropic was this far behind???

Yeah, but actually that's not a good look. Anyone who's used Gemini will know how random it is in terms of getting anything serious done, compared to the rock solid opus experience.

Re: GLM-5.1: Towards Long-Horizon Tasks

#269
post #247

Earlier quoted context omitted.

“GLM5…better than Opus, Codex, Gemini…” What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided. Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?

I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…

FWIW, my experience is the same. Paired with opencode it has been excellent to me.

Re: GLM-5.1: Towards Long-Horizon Tasks

#270
post #254

Earlier quoted context omitted.

Except the rumors are they subsidize even the inference, not that they have capex in training.

The maths shows inference is very profitable. Look at how Google/AWS/Azure change the same rates as Anthropic does for running Claude models.

You're missing the forest for the trees. Per-token pricing is irrelevant when you're just trying to get shit done. I pay 20 bucks a month for OpenAI, but I use likely $200+ a month of tokens just on the coding (and I'm just looking at the raw tokens, this is ignoring all the harnessing on their end). Even OpenAI has said that they're losing money on the 200-dollar subscriptions[1]. This is not a viable business model. Why do you think they are introducing ads this year[2]?

[1] https://fortune.com/2025/01/07/sam-altman-openai-chatgpt-pro...

[2] https://openai.com/index/testing-ads-in-chatgpt/

Post reply on HN