What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
271–280 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#272The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
Distance: 50 meters (very short).
Goal: Get the car washed.
Decision: Drive or walk?
2. Evaluate the Options:
Option A: Walk
Pros: Healthy, environmentally friendly, no need to find parking (though it's close), easy to chat with neighbors.
Cons: The car is left at the car wash. The user can't "walk" the car to the car wash. The car must be moved to the car wash to be washed.
Option B: Drive
Pros: Moves the car to the location where the service is performed. It's the only logical way to get the car washed.
Cons: Driving 50 meters is technically bad for a cold engine (short trips), uses a negligible amount of fuel, might seem lazy.
3. Identify the Logical Flaw/Humor:
The user is asking if they should walk to get their car washed.
Unless the car wash involves someone coming to their house (mobile wash) or carrying the car, the car must be at the car wash.
Walking to the car wash leaves the car at home. Walking back from the car wash leaves the car at the car wash.
Therefore, the user must drive the car to the car wash.
4. Formulate the Answer:
Direct Answer: You should drive.
Reasoning: The car needs to be at the car wash to be washed. You cannot walk the car there.
Nuance/Context: 50 meters is a very short distance. Driving a car such a short distance is generally not great for the engine (doesn't allow oil to circulate/warm up properly), but it is necessary in this specific case.
Alternative (if applicable): Is there a mobile wash option? If so, walking isn't even needed. But assuming a standard car wash bay: Drive.
5. Drafting the response (Internal Monologue/Trial Phrases):
Draft 1: You have to drive. You can't wash the car if it's at your house and you are at the car wash.
Draft 2: You should drive. Since the car needs to be washed, it needs to be at the car wash. Walking there doesn't help get the car clean.
Refining for clarity and wit: Point out the obvious - the car is the object being washed. It needs to move.
Final Polish (similar to the provided good output):
State clearly: You should drive.
Explain why: The car needs to be there.
Add a practical note: Driving 50 meters is a bit silly for engine health, but unavoidable here.
Consider the scenario of "walking back": Maybe the user meant "walk back"? No, the prompt says "drive there or walk there". The intent is transportation to the location.
6. Constructing the final output:
"You should definitely drive."
"Why? Because the car wash cleans the car, so the car needs to be at the car wash. If you walk, the car stays at home."
"Additional thought: Driving 50 meters is technically a 'cold start' short trip which isn't ideal for the engine, but it's the only way to get the service done!"
Final Output Generation (matching the logic above).
You should drive.
Here is why:
1. The car needs to be there: You are getting the car washed, so the car must be at the car wash. You cannot walk the car there by yourself.
2. Efficiency: Driving 50 meters takes about 10 seconds. Walking 50 meters takes about a minute, but you would be leaving the car behind.
The only exception is if you are planning to get a mobile car wash (where they come to you) or if you are dropping the car off and walking back home while they clean it. But to simply "get there," you must drive.
(The above output is from GLM-5)
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#273The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
Maybe, but "intelligence" doesn't have a clear, agreed definition. And calling them "just token generators" skips over how complex that generation actually is.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#274Earlier quoted context omitted.
Sure. My sole point is that calling Opus 4.5 and GPT-5.2 "last generation models" is discounting how good they are. In fact, in my experience, Opus 4.6 isn't much of an improvement over 4.5 for agentic coding. I'm not immediately discounting Z.ai's claims because they showed with GLM-4.7 that they can do quite a lot with very little. And Kimi K2.5 is genuinely a great model, so it's possible for Chinese open-weight m…
I think there are two types of people in these conversations: Those of us who just want to get work done don't care about comparisons to old models, we just want to know what's good right now. Issuing a press release comparing to old models when they had enough time to re-run the benchmarks and update the imagery is a calculated move where they hope readers won't notice. There's another type of discussion where some…
That you think corporations are anything close to quick enough to update their communications on public releases like this only shows that you've never worked in corporate
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#275What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…
> What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips Has any of these outfits ever publicly stated they used Nvidia chips? As in the non-officially obtained 1s. No. > US attempts to contain Chinese AI tech totally failed. Not only that, they cost Nvidia possibly trillions of dollars of exports over the next decade, as the Chinese govt called the American bluff and now active…
Uh yes? Deepseek explicitly said they used H800s [1]. Those were not banned btw, at the time. Then US banned them too. Then US was like 'uhh okay maybe you can have the H200', but then China said not interested.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#276GLM 5 beats Kimi on SWE bench and Terminal bench. If it's anywhere near Kimi in price, this looks great. Edit: Input tokens are twice as expensive. That might be a deal breaker.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#277We already know that intelligence scales with the log of tokens used for reasoning, but Anthropic seems to have much more powerful non-reasoning models than its competitors.
I read somewhere that they have a policy of not advancing capabilities too much, so could it be that they are sandbagging and releasing models with artificially capped reasoning to be at a similar level to their competitors?
How do you read this?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#278Earlier quoted context omitted.
Why do we hear censorship concerns only when it comes Chinese models? Why don't we hear similar stances when Claude or OpenAI releases models? We either set the bar and judge both, or don't complain about censorship
I think more people should spend time talking about this with American models, yeah. If you're interested in that then maybe that can be you. It doesn't have to be the same exact people talking about everything, that's the nice thing about forums. Find your own topic that American models consistently lie or freeze on that Chinese models don't and post about it.
For example,
* I am not expecting Gemini 3 Flash to cure cancer and constantly criticising them for that
* Or I am not expecting Mistral to outcompete OpenAI/Claude on their each release, because talent density and capital is obviously on a different level on OpenAI side
* Or I am not expecting GPT 5.3 saying anytime soon: Yes, Israel committed genocide and politicians covered it up
We should set expectations properly and don't complain about Tianmen every time when Chinese companies are releasing their models and we should learn to appreciate them doing it and creating very good competition and they are very hard working people.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#279Earlier quoted context omitted.
> What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips Has any of these outfits ever publicly stated they used Nvidia chips? As in the non-officially obtained 1s. No. > US attempts to contain Chinese AI tech totally failed. Not only that, they cost Nvidia possibly trillions of dollars of exports over the next decade, as the Chinese govt called the American bluff and now active…
> Has any of these outfits ever publicly stated they used Nvidia chips? As in the non-officially obtained 1s. No. Uh yes? Deepseek explicitly said they used H800s [1]. Those were not banned btw, at the time. Then US banned them too. Then US was like 'uhh okay maybe you can have the H200', but then China said not interested. [1] https://arxiv.org/pdf/2412.19437
Then they haven't. I said the non-officially obtained 1s that they can't / won't mention i.e. those Blackwells etc...
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#280The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
Since the goal is to get your car washed, the car needs to be at the car wash. If you walk, you will arrive at the car wash, but your car will still be sitting at home"
Are you sure that question is from this year?