Earlier quoted context omitted.
This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…
Interesting, what were the instructions if you don't mind sharing?
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
411–420 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#412The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#413It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…
I don't really care about being able to self host these models, but getting to a point where the hosting is commoditised so I know I can switch providers on a whim matters a great deal. Of course, it's nice if I can run it myself as a last resort too.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#414Earlier quoted context omitted.
Your objective has explicit instruction that car has to be present for a wash. Quite a difference from the original phrasing where the model has to figure it out.
That's the answer of his LLM which has decomposed the question and built the answer following the op prompt obviously. I think you didn't get it.
I did get it, and in my view my point still stands. If I need to use special prompts to ask such a simple question, then what are we doing here? The LLMs should be able to figure out a simple contradiction in the question the same way we (humans) do.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#415So that was pony alpha (1). Now what's Aurora Alpha? (1) https://openrouter.ai/openrouter/pony-alpha
It's GPT. Tried and reproduced some polluted single-token Chinese phrases from 4o era.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#416Been using GLM-4.7 for a couple weeks now. Anecdotally, it’s comparable to sonnet, but requires a little bit more instruction and clarity to get things right. For bigger complex changes I still use anthropic’s family, but for very concise and well defined smaller tasks the price of GLM-4.7 is hard to beat.
This aligns very closely with my experience. When left to its own devices, GLM-4.7 frequently tries to build the world. It's also less capable at figuring out stumbling blocks on its own without spiralling. For small, well-defined tasks, it's broadly comparable to Sonnet. Given how incredibly cheap it is, it's useful even as a secondary model.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#417Grey market fast-follow via distillation seems like an inevitable feature of the near to medium future. I've previously doubted that the N-1 or N-2 open weight models will ever be attractive to end users, especially power users. But it now seems that user preferences will be yet another saturated benchmark, that even the N-2 models will fully satisfy. Heck, even my own preferences may be getting saturated already. Op…
In some ways, Opus 4.6 is a step backwards due to massively higher token consumption.
I have started using Gemini Flash on high for general cli questions as I can't tell the difference for those "what's the command again" type questions and it's cheap/fast/accurate.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#418The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
I just copy pasted your question "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" without any further prompt and ran it against GLM 5, GPT 5.2, Opus 4.6, Gemini 3 Pro Preview, through OpenRouter with reasoning effort set to xhigh.
Not a single one said I should walk, they all said to drive.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#419Earlier quoted context omitted.
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
This is really just a meme. People don't know how to use these tools. Here is the response from Gpt-5.2 using my default custom instructions in the mac desktop app. OBJECTIVE: Decide whether to drive or walk to a car wash ~50 meters from home, given typical constraints (car must be present for wash). APPROACH: Use common car-wash workflows + short-distance driving considerations (warm engine, time, parking/queue). No…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#420Earlier quoted context omitted.
If that's what they're tuning for, that's just not what I want. So I'm glad I switched off of Anthropic. What teams of programmers need, when AI tooling is thrown into the mix, is more interaction with the codebase, not less. To build reliable systems the humans involved need to know what was built and how . I'm not looking for full automation, I'm looking for intelligence and augmentation, and I'll give my money and…
I'm not looking for full automation But your boss probably is.