Earlier quoted context omitted.
With Apple devices you get very fast predictions once it gets going but it is inferior to nvidia precisely during prefetch (processing prompt/context) before it really gets going. For our code assistant use cases the local inference on Macs will tend to favor workflows where there is a lot of generation and little reading and this is the opposite of how many of use use Claude Code. Source: I started getting Mac Studi…
> With Apple devices you get very fast predictions once it gets going but it is inferior to nvidia precisely during prefetch (processing prompt/context) before it really gets going I have a Mac and an nVidia build and I’m not disagreeing But nobody is building a useful nVidia LLM box for the price of a $500 Mac Mini You’re also not getting as much RAM as a Mac Studio unless you’re stacking multiple $8,000 nVidia RTX…
GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
421–430 of 540 posts
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#422Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#423Pelican generated via OpenRouter: https://gist.github.com/simonw/cc4ca7815ae82562e89a9fdd99f07... Solid bird, not a great bicycle frame.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#424Earlier quoted context omitted.
> With Apple devices you get very fast predictions once it gets going but it is inferior to nvidia precisely during prefetch (processing prompt/context) before it really gets going I have a Mac and an nVidia build and I’m not disagreeing But nobody is building a useful nVidia LLM box for the price of a $500 Mac Mini You’re also not getting as much RAM as a Mac Studio unless you’re stacking multiple $8,000 nVidia RTX…
Not many are getting useful inference out of a $500 mac mini, due to only having 16GB of RAM.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#425Whoa, I think GPT-5.3-Codex was a disappointment, but GLM-5 is definitely the future!
All I’ve got to add is that GLM-5 is actually just the team at Z.ai getting started. I’m really bullish on this.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#426It's looking like we'll have Chinese OSS to thank for being able to host our own intelligence, free from the whims of proprietary megacorps. I know it doesn't make financial sense to self-host given how cheap OSS inference APIs are now, but it's comforting not being beholden to anyone or requiring a persistent internet connection for on-premise intelligence. Didn't expect to go back to macOS but they're basically the…
> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#427The benchmarks are impressive, but it's comparing to last generation models (Opus 4.5 and GPT-5.2). The competitor models are new, but they would have easily had enough time to re-run the benchmarks and update the press release by now. Although it doesn't really matter much. All of the open weights models lately come with impressive benchmarks but then don't perform as well as expected in actual use. There's clearly…
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was this question: "The car wash is only 50 meters from my house. I want to get my car washed, should I drive there or walk?" Here is Claude's answer just right now: "Walk! At only 50 meters…
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#428Earlier quoted context omitted.
It's interesting how some features, such as green grass, a blue sky, clouds, and the sun, are ubiquitous among all of these models' responses.
If you were a pelican, wouldn't you want to go cycling on a sunny day? Do electric pelicans dream of touching electric grass?
That would be shocking news to me.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#429Really impressive benchmarks. It was commonly stated that open source models were lagging 6 months behind state of the art, but they are likely even closer now.
Open-weights models are still lagging quite a bit behind SOTA. E.g. there's still no open model that can match GPT-5 Pro or Gemini 2.5 Pro, and the latter is almost a year old by now.
Re: GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
#430Earlier quoted context omitted.
> doesn't make financial sense to self-host I guess that's debatable. I regularly run out of quota on my claude max subscription. When that happens, I can sort of kind of get by with my modest setup (2x RTX3090) and quantized Qwen3. And this does not even account for privacy and availability. I'm in Canada, and as the US is slowly consumed by its spiral of self-destruction, I fully expect at some point a digital iron…
Anthropic has very tight limits, so you're basically using the worst (pricing-wise) SOTA cloud model as your baseline. I have $200 subs for both Claude and OpenAI, and I also bump into limits with Claude all the time, whether coding or research. With Codex, I ran into the limit once so far, and that's in a month of very heavy (sometimes literally 24 hours around the clock, leaving long-running tasks overnight) use.
I hope too many of us won't be doing this and cause Google to add limits! My hope is Google sees the benefit in this and goes all in - continues to let people decide which Google hosted model to use, including their own.