Earlier quoted context omitted.
I mean it's not hard to understand that if good model can run on consumer hardware, even better models can run in data centers
Larger, yes, absolutely. Better? Right now it seems that bigger is better, but if we are thinking about long term future, it's not obvious that there isn't a point of diminishing returns with regards to size. I can also imagine a breakthrough, where models become much smaller, with the same or better capabilities as the current, very large ones.
GLM-5.1: Towards Long-Horizon Tasks
211–220 of 285 posts
Re: GLM-5.1: Towards Long-Horizon Tasks
#212Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…
I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…
Re: GLM-5.1: Towards Long-Horizon Tasks
#213We're still adding samples, but some early takeaways from benchmarking on https://gertlabs.com : Contrary to the model card, its one-shot performance is more impressive than its agentic abilities. On both metrics, GLM 5.1 is competitive with frontier models. But keeping in mind this is an open source model operating near the frontier, it's nothing short of incredible. I suspect 2 issues with the model are keeping it…
Re: GLM-5.1: Towards Long-Horizon Tasks
#214Earlier quoted context omitted.
as a local LLM novice, do you have any recommended reading to bootstrap me on selecting hardware? It has been quite confusing bring a latecomer to this game. Googling yields me a lot of outdated info.
First answer: If you haven't, give it a shot on whatever you already have. MoE models like Qwen3 and GPT-OSS are good on low-end hardware. My RTX 4060 can run qwen3:30b at a comfortable reading pace even though 2/3 of it spills over into system RAM. Even on an 8-year-old tiny PC with 32gb it's still usable. Second answer: ask an AI, but prices have risen dramatically since their training cutoff, so be sure to get the…
I run qwen 122b with Claude code and nanoclaw, it's pretty decent but this stuff is nowhere prime time ready, but super fun to tinker with. I have to keep updating drivers and see speed increases and stability being worked on. I can even run much larger models with llama.cpp (--fit on) like qwen 397b and I suppose any larger model like GLM, it's slow but smart.
Re: GLM-5.1: Towards Long-Horizon Tasks
#215Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…
I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…
Re: GLM-5.1: Towards Long-Horizon Tasks
#216Earlier quoted context omitted.
> Top-tier models will never run on desktop machines Sorry, but you don't know that
I mean it's not hard to understand that if good model can run on consumer hardware, even better models can run in data centers
Re: GLM-5.1: Towards Long-Horizon Tasks
#217Earlier quoted context omitted.
I'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long ho…
Does anyone have inside info on what these Huawai chips look like? I know Google has a Torus architecture unlike Nvidias fully connected one. Maybe it’s a similar architectural decision on the huawai chips that leads to bottlenecks in serving?
>For AI computing, the Atlas 950 SuperPoD, powered by UnifiedBus, integrates 64 NPUs per cabinet and can scale up to 8,192 NPUs, delivering superior performance for large-scale AI training and high-concurrency inference.
Re: GLM-5.1: Towards Long-Horizon Tasks
#218Z.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not…
https://huggingface.co/trohrbaugh/gemma-4-31b-it-heretic-ara...
which was produced immediately after Google released their new Gemma 4 model.
Re: GLM-5.1: Towards Long-Horizon Tasks
#219Earlier quoted context omitted.
Like computing used to be. When I first compiled a Linux kernel it ran overnight on a Pentium-S. I had little idea what I was doing, probably compiled all the modules by mistake.
At least the compiler was free
Re: GLM-5.1: Towards Long-Horizon Tasks
#220Earlier quoted context omitted.
This has got to be bait.. 1) OpenAI and Anthropic are killing it, and continue to do so, their coding tools are unmatched for professionals. 2) Local models don't hold a candle to SOTA models and there's nothing on the horizon that indicates that consumers will be able to run anything close to what you can get in a data center. 3) Coding is a killer product, OpenAI and Anthropic are raking in the cash. The top 3 apps…
I don't want to respond to 100 comments about the same thing, and this one happens to be on top, so, in my humble opinion: (1): You don't have to be an Ed Zitron disciple to infer that OpenAI and Anthropic are likely overvalued and that Nvidia is selling everyone shovels in a gold rush. AI is a game-changing technology, but a shitty chat interface does not a company make. OpenAI and Anthropic need to recoup astronomi…
Anthropic is up to $30B annual recurring revenue. I wish I had failing business models like that.
> Token prices are significantly subsidized and anyone that does any serious work with AI can tell you this. Go use an almost-SOTA model (a big Deepseek or Qwen model) offered by many bare-metal providers and you'll see what "true" token prices should look like.
I'm not sure what think you are saying here, but if you look at the providers for both "almost-SOTA model (a big Deepseek or Qwen model)" or at the price for Claude on AWS Bedrock, Azure or on GCP you will quickly see inference is very profitable.