GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
361–370 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#362> GLM-5 can turn text or source materials directly into .docx, .pdf, and .xlsx files—PRDs, lesson plans, exams, spreadsheets, financial reports, run sheets, menus, and more.
A new type of model has joined the series, GLM-5-Coder.
GLM-5 was trained on Huawei Ascend, last time when DeepSeek tried to use this chip, it flopped and they resorted to Nvidia again. This time seems like a success.
Looks like they also released their own agentic IDE, https://zcode.z.ai
I don’t know if anyone else knows this but Z.ai also released new tools excluding the Chat! There’s Zread (https://zread.ai), OCR (seems new? https://ocr.z.ai), GLM-Image gen https://image.z.ai and Voice cloning https://audio.z.ai
If you go to chat.z.ai, there is a new toggle in the prompt field, you can now toggle between chat/agentic. It is only visible when you switch to GLM-5.
Very fascinating stuff!
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#363What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#364Earlier quoted context omitted.
I think you're seriously underestimating how much effort the fine tuning at their scale takes and what impact it has. They don't pack every edge case into the system prompt either. It's not like they update the model every few hours or even care about memes. If they seriously did, they'd force-delegate spelling questions to tool calls.
Could it be the model is constantly searching its own name for memes, or checking common places like HN and updating accordingly? I have no idea how real-time these things are, just asking.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#365Earlier quoted context omitted.
Thank you for continuing to maintain the only benchmarking system that matters! Context for the unaware: https://simonwillison.net/tags/pelican-riding-a-bicycle/
They will start to max this benchmark as well at some point.
It's just an experiment on how different models interpret a vague prompt. "Generate an SVG of a pelican riding a bicycle" is loaded with ambiguity. It's practically designed to generate 'interesting' results because the prompt is not specific.
It also happens to be an example of the least practical way to engage with an LLM. It's no more capable of reading your mind than anyone or anything else.
I argue that, in the service of AI, there is a lot of flexibility being created around the scientific method.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#366The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
What a strangely hostile statement on an open weight model. Running like 20 benchmark evaluations isn't trivial by itself, and even updating visuals and press statements can take a few days at a tech company. It's literally been 5 days since this "new generation" of models released. GPT-5.3(-codex) can't even be called via API, so it's impossible to test for some benchmarks. I notice the people who endlessly praise c…
open weight models are not there at all yet.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#367Earlier quoted context omitted.
If you're asking simple riddles, you shouldn't be paying for SOTA frontier models with long context. This is a silly test for the big coding models. This is like saying "all calculators are the same, nobody needs a TI-89!" and then adding 1+2 on a pocket calculator to prove your point.
No it’s like having a calculator which is unable to perform simple arithmetic, but lots of people think it is amazing and sentient and want to talk about that instead of why it can’t add 2 + 2.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#368Earlier quoted context omitted.
They will start to max this benchmark as well at some point.
It's not a benchmark though, right? Because there's no control group or reference. It's just an experiment on how different models interpret a vague prompt. "Generate an SVG of a pelican riding a bicycle" is loaded with ambiguity. It's practically designed to generate 'interesting' results because the prompt is not specific. It also happens to be an example of the least practical way to engage with an LLM. It's no mo…
For the last generation of models, and for today's flash/mini models, I think there is still a not-unreasonable binary question ("is this a pelican on a bicycle?") that you can answer by just looking at the result: https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#369Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#370Earlier quoted context omitted.
They will start to max this benchmark as well at some point.
It's not a benchmark though, right? Because there's no control group or reference. It's just an experiment on how different models interpret a vague prompt. "Generate an SVG of a pelican riding a bicycle" is loaded with ambiguity. It's practically designed to generate 'interesting' results because the prompt is not specific. It also happens to be an example of the least practical way to engage with an LLM. It's no mo…