Live data from Hacker News

GLM-5.1: Towards Long-Horizon Tasks

z.ai

121–130 of 285 posts

Re: GLM-5.1: Towards Long-Horizon Tasks

#121
post #65

Comments here seem to be talking like they've used this model for longer than a few hours -- is this true, or are y'all just sharing your initial thoughts?

My local tennis court's reservation website was broken and I couldn't cancel a reservation, and I asked GLM-5.1 if it can figure out the API. Five minutes later, I check and it had found a /cancel.php URL that accepted an ID but the ID wasn't exposed anywhere, so it found and was exploiting a blind SQL injection vulnerability to find my reservation ID. Overeager, but I was really really impressed.

> Five minutes later, I check and it had found a /cancel.php URL that accepted an ID but the ID wasn't exposed anywhere, so it found and was exploiting a blind SQL injection vulnerability to find my reservation ID.

xkcd was prescient once again... https://xkcd.com/416/

Re: GLM-5.1: Towards Long-Horizon Tasks

#122
This is the flip side of the Project Glasswing stuff...

Everyone else isn't that far behind and they aren't all gonna just wall off their new model.

A reason that Anthropic will eventually give is 'the competition can do what Glasswing can do so what's the point limiting it'.

Re: GLM-5.1: Towards Long-Horizon Tasks

#123
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

I don't see how its possible to think this. AI coding assistants are some of the most useful technologies ever created, and model quality is by far the most important thing, so I doesn't make sense why local inference would be the path forward unless something fundamentally changes about hardware.

Re: GLM-5.1: Towards Long-Horizon Tasks

#125
post #124

GLM 5.1 does worse than GLM 5 in my tests[0] (both medium reasoning OR no reasoning). I think the model is now tuned more towards agentic use/coding than general intelligence. [0]: https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-...

The (none) version especially shows considerable degradation.

Re: GLM-5.1: Towards Long-Horizon Tasks

#128
post #121
post #65

Earlier quoted context omitted.

My local tennis court's reservation website was broken and I couldn't cancel a reservation, and I asked GLM-5.1 if it can figure out the API. Five minutes later, I check and it had found a /cancel.php URL that accepted an ID but the ID wasn't exposed anywhere, so it found and was exploiting a blind SQL injection vulnerability to find my reservation ID. Overeager, but I was really really impressed.

> Five minutes later, I check and it had found a /cancel.php URL that accepted an ID but the ID wasn't exposed anywhere, so it found and was exploiting a blind SQL injection vulnerability to find my reservation ID. xkcd was prescient once again... https://xkcd.com/416/

Hell, this one time, my AI assistant hacked itself trying to book an appointment for me!

Re: GLM-5.1: Towards Long-Horizon Tasks

#129
post #119
post #118

Every single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)

What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?

Is it so hard to project out a couple product cycles? Computers get better. We’ve gone from $50k workstation to commodity hardware before several times
Post reply on HN