The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
341–350 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#342Earlier quoted context omitted.
This Pelican benchmark has become irrelevant. SVG is already ubiquitous. We need a new, authentic scenario.
Like identifying names of skateboard tricks from the description? https://skatebench.t3.gg/
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#343Earlier quoted context omitted.
It's unclear where the car is currently from your phrasing. If you add that the car is in your garage, it says you'll need to drive to get the car into the wash.
Do you think the average person would need this sort of clarification? How many of us would have recommended to walk?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#344Earlier quoted context omitted.
US Secretary of State Bressent just publicly said that the US needs to get along and cooperate with China. His tone was so different than previously in the last year that I listened to the video clip twice. Obviously for the average US tax payer getting along with China is in our interests - not so much our economic elites. I use both Chinese and US models, and Mistral in Proton’s private chat. I think it makes sense…
>His tone was so different than previously in the last year that I listened to the video clip twice. US bluff got called. A year back it looked like US held all the cards and could squeeze others without negative consequences. i.e. have cake and eat it too Since then: China has not backed down, Europe is talking de-dollarization, BRICS is starting to find a new gear on separate financial system, merciless mocking acr…
And yes, the consequence is strengthening the actual enemies of the USA, their AI progress is just one symptom of this disastrous US administration and the incompetence of Donald Trump. He really is the worst President of the USA ever, even if you were to just judge him on his leadership regarding technology... and I'm saying this while he is giving a speech about his "clean beautiful coal" right now in the White House.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#345The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
"" [...] Since you need to get your car washed, you have to bring the car to the car wash—walking there without the vehicle won't accomplish your goal [...] If it's a self-service wash, you could theoretically push the car 50 meters if it's safe and flat (unusual, but possible) [..] Consider whether you really need that specific car wash, or if a mobile detailing service might come to you [...] """
Which seems slightly (unintentionally) funny.
But to be fair all the Gemini (including flash) and GPT models I tried did understand the quesiton.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#346Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#347I got fed up with GLM-4.7 after using it for a few weeks; it was slow through z.ai and not as good as the benchmarks lead me to believe (esp. with regards to instruction following) but I'm willing to give it another try.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#348I got fed up with GLM-4.7 after using it for a few weeks; it was slow through z.ai and not as good as the benchmarks lead me to believe (esp. with regards to instruction following) but I'm willing to give it another try.
Try Cerberas
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#349Earlier quoted context omitted.
Thank you for continuing to maintain the only benchmarking system that matters! Context for the unaware: https://simonwillison.net/tags/pelican-riding-a-bicycle/
It's interesting how some features, such as green grass, a blue sky, clouds, and the sun, are ubiquitous among all of these models' responses.
Do electric pelicans dream of touching electric grass?
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#350The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
I think the only advantage that closed models have are the tools around them (claude code and codex). At this point if forced I could totally live with open models only if needed.