Live data from Hacker News

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

arxiv.org

21–30 of 35 posts

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#21
post #10

Click coordinates. Agentic GUI is really annoying when the multi-modal agent cannot click on x,y coordinates. I tested Qwen3.6, Gemma4, Nemotron3-nano-omni. They fully hallucinate x,y coords. (did not try GLM-5V yet) GPT-5.5 can easily do it. But also Vocaela, a tiny 500M model, is quite good at it. Hope they improve the training for x,y clicking soon on the smallish multi-modals. Recently slopped a http service toge…

I've had lots of success with generating coordinates and answering questions using the UI-TARS model https://github.com/bytedance/UI-TARS.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#22
post #10

Click coordinates. Agentic GUI is really annoying when the multi-modal agent cannot click on x,y coordinates. I tested Qwen3.6, Gemma4, Nemotron3-nano-omni. They fully hallucinate x,y coords. (did not try GLM-5V yet) GPT-5.5 can easily do it. But also Vocaela, a tiny 500M model, is quite good at it. Hope they improve the training for x,y clicking soon on the smallish multi-modals. Recently slopped a http service toge…

Qwen3.5 is able to output click coordinates and bounding boxes just fine, as values normalized to 0..1000, I’d hope Qwen3.6 didn’t loose this capability.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#24
post #20
post #14

We just migrated an AI agent from Kimi to GLM and frankly I am surprised by the results. It feels premium. However, both Kimi and GLM can end up in doom loops so be careful how you use them. Without a proper harness the agent can easily get into some tricky situations with no escape. We had to develop new heuristics in our cloud harness just because of this but I am really grateful that we did as the platform feels n…

What version of Kimi was that using? Do you have any specific insight on Kimi vs GLM in real world scenarios?

[dead]

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#26
post #10

Click coordinates. Agentic GUI is really annoying when the multi-modal agent cannot click on x,y coordinates. I tested Qwen3.6, Gemma4, Nemotron3-nano-omni. They fully hallucinate x,y coords. (did not try GLM-5V yet) GPT-5.5 can easily do it. But also Vocaela, a tiny 500M model, is quite good at it. Hope they improve the training for x,y clicking soon on the smallish multi-modals. Recently slopped a http service toge…

I've had lots of success with generating coordinates and answering questions using the UI-TARS model https://github.com/bytedance/UI-TARS .

I’d also checkout midscene, you can set the model and UI-TARS works but you can also use qwen vision models and it works.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#27
post #3

z.ai will use quantized models in off hours. Buyer beware

I hear a lot of people complaining, I am on their Max plan, I never hit limits, use it non-stop and overall it has been fantastic experience.

Has 5.1 reliability improved? I would love to use it again. The inference was just too unreliable when it was first released.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#28
post #3

z.ai will use quantized models in off hours. Buyer beware

I hear a lot of people complaining, I am on their Max plan, I never hit limits, use it non-stop and overall it has been fantastic experience.

Same feeling here on the Pro plan. I’m still in the old plan without the Weekly quotas, but never never exhausted the 5H so far.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#29
post #20
post #14

We just migrated an AI agent from Kimi to GLM and frankly I am surprised by the results. It feels premium. However, both Kimi and GLM can end up in doom loops so be careful how you use them. Without a proper harness the agent can easily get into some tricky situations with no escape. We had to develop new heuristics in our cloud harness just because of this but I am really grateful that we did as the platform feels n…

What version of Kimi was that using? Do you have any specific insight on Kimi vs GLM in real world scenarios?

K2.6. My experience is anecdotal but GLM was able to complete the task more throughly and quicker.

Re: GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

#30

I've been using GLM pretty much exclusively last 6-8 months. I have access to Anthropic and OpenAI models and others. I always keep returning to GLM, it isn't the best, sometimes I would go to Codex to help it, but overall, especially with Turbo, it is everyday good model. Turbo makes a huge difference in everyday use because it saves you time and you are not in the mood always to wait endlessly.

> I've been using GLM pretty much exclusively last 6-8 months. I have access to Anthropic and OpenAI models and others. I always keep returning to GLM, it isn't the best

Very interesting. What sort of tasks do you use it for, and what client do you use?

When you want to use a custom client and a coding plan to control costs (daily use, a few hundred USD/m budget), this is the landscape:

- Anthropic/Google: Deterring custom clients actively

- OpenAI: Grey area.

- Z.ai: Technically only allows clients in their (large) approved list of clients. Likely won't actively ban custom clients.

- Moonshot: Seem to allow custom clients?

- DeepSeek/Alibaba: No coding plans at this time

Post reply on HN