Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly some benchmaxxing going on.
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
181–190 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#182Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#183Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#184Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#185Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#186why don't they publish at ARC-AGI ? too expensive?
Arc agi was never a good benchmark that tested spatial understanding more than reasoning. I'm glad it's no longer popular
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#187In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general.
To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#188Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#189Earlier quoted context omitted.
Your $5,000 PC with 2 GPUs could have bought you 2 years of Claude Max, a model much more powerful and with longer context. In 2 years you could make that investment back in pay raise.
> In 2 years you could make that investment back in pay raise. you can't be a happy uber driver making more money in the next 24 months by having a fancy car fitted with the best FSD in town when all cars in your town have the same FSD.