Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
How many pelican riding bicycle SVGs were there before this test existed? What if the training data is being polluted with all these wonky results...
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
321–330 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#322Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#323Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
This Pelican benchmark has become irrelevant. SVG is already ubiquitous. We need a new, authentic scenario.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#324What is truly amazing here is the fact that they trained this entirely on Huawei Ascend chips per reporting [1]. Hence we can conclude the semiconductor to model Chinese tech stack is only 3 months behind the US, considering Opus 4.5 released in November. (Excluding the lithography equipment here, as SMIC still uses older ASML DUV machines) This is huge especially since just a few months ago it was reported that Deep…
US Secretary of State Bressent just publicly said that the US needs to get along and cooperate with China. His tone was so different than previously in the last year that I listened to the video clip twice. Obviously for the average US tax payer getting along with China is in our interests - not so much our economic elites. I use both Chinese and US models, and Mistral in Proton’s private chat. I think it makes sense…
US bluff got called. A year back it looked like US held all the cards and could squeeze others without negative consequences. i.e. have cake and eat it too
Since then: China has not backed down, Europe is talking de-dollarization, BRICS is starting to find a new gear on separate financial system, merciless mocking across the board, zero progress on ukraine, fed wobbled, focus on gold as alternate to US fiat, nato wobbled, endless scandals, reputation for TACO, weak employment, tariff chaos, calls for withdrawal of gold from US's safekeeping, chatter about dumping US bonds, multiple major countries being quite explicit about telling trump to get fucked
Not at all surprised there is a more modest tone...none of this is going the "without negative consequences" way
>Mistral in Proton’s private chat
TIL
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#325Earlier quoted context omitted.
Sure. My sole point is that calling Opus 4.5 and GPT-5.2 "last generation models" is discounting how good they are. In fact, in my experience, Opus 4.6 isn't much of an improvement over 4.5 for agentic coding. I'm not immediately discounting Z.ai's claims because they showed with GLM-4.7 that they can do quite a lot with very little. And Kimi K2.5 is genuinely a great model, so it's possible for Chinese open-weight m…
I think there are two types of people in these conversations: Those of us who just want to get work done don't care about comparisons to old models, we just want to know what's good right now. Issuing a press release comparing to old models when they had enough time to re-run the benchmarks and update the imagery is a calculated move where they hope readers won't notice. There's another type of discussion where some…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#326Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
Thank you for continuing to maintain the only benchmarking system that matters! Context for the unaware: https://simonwillison.net/tags/pelican-riding-a-bicycle/
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#327Earlier quoted context omitted.
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
It's unclear where the car is currently from your phrasing. If you add that the car is in your garage, it says you'll need to drive to get the car into the wash.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#328Earlier quoted context omitted.
I just ran this with Gemini 3 Pro, Opus 4.6, and Grok 4 (the models I personally find the smartest for my work). All three answered correctly.
They had plenty of time to update their system prompts so they don't be embarrassed. I noticed whenever such meme comes out, if you check immediately you can reproduce it yourself, but after a free hours it's already updated.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#329Certainly seems to remember things better and is more stable on long running tasks.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#330The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
Anthropic, OpenAI and Google have real user data that they can use to influence their models. Chinese labs have benchmarks. Once you realize this, it's obvious why this is the case. You can have self-hosted models. You can have models that improve based on your needs. You can't have both.