Earlier quoted context omitted.
Yeah it seems they did not align it to much, at least for now. Yesterday it helped me bypass the bot detection on a local marketplace. that i wanted to scrap some listing for my personal alerting system. Al the others failed but glm5.1 found a set of parameters and tweaks how to make my browser in container not be detected.
Model doing what the user wants with high quality is definitely aligned in my book.
GLM-5.1: Towards Long-Horizon Tasks
261–270 of 285 posts
Re: GLM-5.1: Towards Long-Horizon Tasks
#262Earlier quoted context omitted.
Computers get better and cheaper. That’s not a forever problem.
Source? GPU and RAM prices have definitely not made consumer PC's cheaper than they were before bitcoin blew up or before AI blew up. Maybe you could make an argument that they are more cost efficient for the price point... But that's not the same as cheaper when every application or program is poorly optimized. For example why would a browser take up more than a GB or two of RAM? And I'd postulate that R&D to develo…
We've had RAM shocks before. We nerds can't control the Wall Street or Virginians who like to break the world every so often for the lulz. However, a wobble on the curve doesn't change the curve's destination.
Re: GLM-5.1: Towards Long-Horizon Tasks
#263Earlier quoted context omitted.
Computers get better and cheaper. That’s not a forever problem.
Source? GPU and RAM prices have definitely not made consumer PC's cheaper than they were before bitcoin blew up or before AI blew up. Maybe you could make an argument that they are more cost efficient for the price point... But that's not the same as cheaper when every application or program is poorly optimized. For example why would a browser take up more than a GB or two of RAM? And I'd postulate that R&D to develo…
Re: GLM-5.1: Towards Long-Horizon Tasks
#264Earlier quoted context omitted.
The 4-bit quants are 350GB, what hardware are you talking about?
qwen3:0.6b is 523mb, what model are you talking about? You seem to have a specific one in mind but the parent comment doesn't mention any. For a hobby/enthusiast product, and even for some useful local tasks, MoE models run fine on gaming PCs or even older midrange PCs. For dedicated AI hardware I was thinking of Strix Halo - with 128gb is currently $2-3k. None of this will replace a Claude subscription.
1) What are you going to use that for? 0.6 model gives you what you could get from Siri when it first launched at most unless you do some tunning.
2) Pretty clear that they are talking about GLM-5.1 4-bit quant.
Re: GLM-5.1: Towards Long-Horizon Tasks
#265Earlier quoted context omitted.
“GLM5…better than Opus, Codex, Gemini…” What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided. Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?
I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…
You’re making a claim, and I’m pointing out that it’s unsubstantiated and not consistent with any other source of data, including that internal to the company that makes the model.
I hope you can see that that’s different than saying it’s worked well for me
Re: GLM-5.1: Towards Long-Horizon Tasks
#266Earlier quoted context omitted.
If it's relevant to the discussion, I hope not. I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation. Let me know if you know of a better platform you can use…
It's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu...
Who knew Anthropic was this far behind???
Re: GLM-5.1: Towards Long-Horizon Tasks
#267Earlier quoted context omitted.
I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…
It’s not a flame war, and you’re not just sharing your experience and encouraging others to try it out. You’re making a claim, and I’m pointing out that it’s unsubstantiated and not consistent with any other source of data, including that internal to the company that makes the model. I hope you can see that that’s different than saying it’s worked well for me
I do not think that anyone who read my comment understood it differently. But I grant you this point, this is just my opinion based on my personal experience not the result of a scientific study.
Once this is said, i wasn't submitting a scientific paper for preprint, just posting my opinion on an internet forum.
Not sure why you are making such a big deal out of it, especially for something for which people can decide within minutes if it works for them or not. And I haven't seen you nitpick on other people saying that all Chinese models are garbage incapable of doing even the most basic task, without quoting any study. This kind of scrutiny tends to be one-sided.
Edit: and regarding what the z.ai team is saying about their models, just check their Discord and the articles they link there. They themselves say that their latest models have leading performance on a number of aspects. It is misleading to suggest that the authors of the model are not proudly saying that their models have best in class performance.
Re: GLM-5.1: Towards Long-Horizon Tasks
#268Earlier quoted context omitted.
It's a great benchmark. Don't listen to the haters. This one is especially interesting. https://aibenchy.com/compare/anthropic-claude-sonnet-4-6-med...
This one's even more interesting https://aibenchy.com/compare/anthropic-claude-opus-4-6-mediu... Who knew Anthropic was this far behind???
Re: GLM-5.1: Towards Long-Horizon Tasks
#269Earlier quoted context omitted.
“GLM5…better than Opus, Codex, Gemini…” What wild claim to make. Unsupported by benchmarks, unsupported by the consensus of the community, no evidence provided. Sounds like in another comment here even the GLM5 team concedes they are behind the frontier wrt tool calling, do you know something they don’t?
I know my use case and my personal experience :) i am not trying to pretend that it is the best in benchmarks, just sharing my experience so people know that some folks are having a very good experience with GLM models, compared to the competition. My only goal is to encourage people to try it out so they can see if it moves the needle for them, because there are fair chances that it will. I am not trying to start a…
Re: GLM-5.1: Towards Long-Horizon Tasks
#270Earlier quoted context omitted.
Except the rumors are they subsidize even the inference, not that they have capex in training.
The maths shows inference is very profitable. Look at how Google/AWS/Azure change the same rates as Anthropic does for running Claude models.
[1] https://fortune.com/2025/01/07/sam-altman-openai-chatgpt-pro...