GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
371–380 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#372Earlier quoted context omitted.
This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…
Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.
Which is exactly how you're supposed to prompt an LLM, is the fact that giving a vague prompt gives poor results really suprising?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#373Earlier quoted context omitted.
This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…
Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#374Earlier quoted context omitted.
Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.
> Your objective has explicit instruction that car has to be present for a wash. Which is exactly how you're supposed to prompt an LLM, is the fact that giving a vague prompt gives poor results really suprising?
The whole idea of this question is to show that pretty often implicit assumptions are not discovered by the LLM.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#375Earlier quoted context omitted.
Anthropic, OpenAI and Google have real user data that they can use to influence their models. Chinese labs have benchmarks. Once you realize this, it's obvious why this is the case. You can have self-hosted models. You can have models that improve based on your needs. You can't have both.
zAI, minimax and Kimi have plenty of subscriber usage on their own platforms. They get real data just as well. Less or it maybe but it's there.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#376Earlier quoted context omitted.
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#377While GLM-5 seems impressive, this release also included lots of new cool stuff! > GLM-5 can turn text or source materials directly into .docx, .pdf, and .xlsx files—PRDs, lesson plans, exams, spreadsheets, financial reports, run sheets, menus, and more. A new type of model has joined the series, GLM-5-Coder. GLM-5 was trained on Huawei Ascend, last time when DeepSeek tried to use this chip, it flopped and they resor…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#378Earlier quoted context omitted.
I just ran this with Gemini 3 Pro, Opus 4.6, and Grok 4 (the models I personally find the smartest for my work). All three answered correctly.
They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#379Earlier quoted context omitted.
zAI, minimax and Kimi have plenty of subscriber usage on their own platforms. They get real data just as well. Less or it maybe but it's there.
I'm going to claim that the majority of those users are optimizing for cost and not correctness and therefore the quality of data collected from those sessions is questionable. If you're working on something of consequence, you're not using those platforms. If you're a tinkerer pinching pennies, sure.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#380What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…
I've only seen information suggesting that you can run inference with Ascends, which is obviously a very different thing. The source you link also just says: "The latest model was developed using domestically manufactured chips for inference, including Huawei's flagship Ascend chip and products from leading industry players such as Moore Threads, Cambricon and Kunlunxin, according to the statement."