I've got a 48GB M5 - What's the best I can run on that atm (with a bit of headroom).
Show HN: Getting GLM 5.2 running on my slow computer
191–200 of 269 posts
Re: Show HN: Getting GLM 5.2 running on my slow computer
#192Re: Show HN: Getting GLM 5.2 running on my slow computer
#193Excuse my ignorance. Could one just say, "One expert is all I can handle" and strip the others from the model?
Re: Show HN: Getting GLM 5.2 running on my slow computer
#194AMD Ryzen Threadripper PRO 5975WX — 32 cores / 64 threads, Zen3 (znver3), AVX2+FMA (no AVX-512/VNNI), 128GB RAM, Kingston SKC3000D 4TB NVMe (PCIe4). Disk gets around 7GB/s. It took a little tuning (for example pinning to 32 physical cores instead of the 64 threads), but with that and --topp 0.7, got 0.44 tok/s on a cold start. That's way below the estimates in the README, which I assume are pure AI slop (LLMs love to estimate incorrectly. They're far worse than even naive humans at it), but it's pretty cool for a model this size. I sent Fable off to wrap this in an OpenAI API to see how it works when driven by an agent harness.
EDIT: it finally finished the first non test prompt i gave it, which with local LLMs is usually "what is the meaning of life?" (who knows, maybe one of them will finally answer). It got stuck in a loop, which is not encouraging, so there's a lot of work to do to make this a viable local coding tool:
> The meaning of life is one of the oldest and greatest questions in human history, yet strangely, there is no single, universally agreed-upon answer. Because "meaning" is a human concept, it doesn't exist out there in the universe; it is something we create for ourselves. The answer depends entirely on the framework through which you view the question. Here are the most common ways to answer it. The meaning of life is the meaning you give to it. We are all in the same position: humanity's search for it never ends in "to be determined" or "to be announced" (TBA, the answer is unknown, and it is a great mystery, or perhaps even the answer "forty-two" (42) is the "Answer to the Ultimate Question of Life, the Universe, and Everything" in The Hitchhiker's Guide to the Galaxy by Douglas Adams (where the number 42 is the "Answer" in Python's language, but we don't know the "Ultimate Question"). Here is a joke that works under the frame of "A..." (any answer): "A clean desk is a..." (42 is a "portmanteau" of words and just a great big "Ad..." (Ad-100) and "A&d" (100)). Life is a deep and strange and we search for meaning in it. "I think, therefore,..." (Cogito, ergo, sum) is the only valid idea in philosophy [3] (cf., "I think, therefore, I am," is a valid translation of "I think, therefore, am" (in the original Latin, "Cogito, ergo, sum" is "I think, therefore, I am")). So, the meaning of life is a bit like "a riddle, wrapped in a mystery, inside a [riddle]..." (G. K. Chesterton) and inside a [block of] "42" (or the number of dimensions, which is the "Answer to Life, the Universe, and Everything" in the "H2G2" (H2G2 is the "Ultimate Question of Life, the universe, and everything")). The "H2G2" is a "puzzle, wrapped in a mystery, inside an enigma" (cf. [3]). We are all in the same position, but we all have to give it a meaning. our own meaning. The meaning of life is what you make of it. The meaning of life is to live for the greater good. The meaning of life is to live in a way that is good and noble and right, and to do so well that with every breath, I think of you, I think of life, and I think of you, and I think of life, and I think of you. (cf. [3]) If life in the universe is a "great question," the answer is 42. The meaning of life is the meaning you make it. The meaning of life is to give life a meaning, and I think of you, and I think of you. So, the answer to the ultimate question of life, the universe, and everything is: 42. The meaning of life is 42. The meaning of life is the meaning of life. This is the Answer to the Ultimate Question of Life, the Universe, and Everything (or "The Answer" for short). It is the Answer to "the" Ultimate Question of Life, the Universe, and Everything. (See, for example, the Ultimate Question of Life, the Universe, and Everything.) This is the answer to the Ultimate Question. This is the Answer. (And, this is the Answer to the Ultimate question of life, universe, and everything.) The meaning of life is the meaning you give to it. The meaning of life is to give it a meaning. The meaning of life is the meaning you give it. The meaning of life is the meaning of life. The meaning of life is the meaning of life. The meaning of life is the meaning of life. (This is a list of the possible meanings of the universe of life. It's a list of the most common and accepted answers. "What is the meaning of life?" The answer is 42. The meaning of life is the meaning of life. The answer is 42.) (See also: [3] for a list of possible meanings.) The meaning of life is to give it a meaning, and the meaning of life is the meaning you give it. The meaning of life is the meaning of life. The meaning of life is the meaning of life. The meaning of life is the meaning of life. (This is the answer to the Ultimate Question of life, the universe, and everything.) (This is the answer to the Ultimate Question.) The meaning of life is 42. The meaning of life is 42. The meaning of life is 42. (See also: [3]) (The answer to the Ultimat
Re: Show HN: Getting GLM 5.2 running on my slow computer
#195Tried this out on my ThreadRipper: AMD Ryzen Threadripper PRO 5975WX — 32 cores / 64 threads, Zen3 (znver3), AVX2+FMA (no AVX-512/VNNI), 128GB RAM, Kingston SKC3000D 4TB NVMe (PCIe4). Disk gets around 7GB/s. It took a little tuning (for example pinning to 32 physical cores instead of the 64 threads), but with that and --topp 0.7, got 0.44 tok/s on a cold start. That's way below the estimates in the README, which I as…
Re: Show HN: Getting GLM 5.2 running on my slow computer
#196Tried this out on my ThreadRipper: AMD Ryzen Threadripper PRO 5975WX — 32 cores / 64 threads, Zen3 (znver3), AVX2+FMA (no AVX-512/VNNI), 128GB RAM, Kingston SKC3000D 4TB NVMe (PCIe4). Disk gets around 7GB/s. It took a little tuning (for example pinning to 32 physical cores instead of the 64 threads), but with that and --topp 0.7, got 0.44 tok/s on a cold start. That's way below the estimates in the README, which I as…
I’ve been looking at exactly this kind of system (in a Lenovo P620) to fill with external GPUs (powered externally and ribbon cabled into the pcie slots). What would you say was your best performing model on this system? And do you get any useful work done with it or are you still dependent on SOTA models online?
Re: Show HN: Getting GLM 5.2 running on my slow computer
#197My main question is whether when put into practical use, this can be measured in tokens/second, or more like 1 token per minute... I have seen locally hosted LLM that are as slow as 1 tok/second still be very useful if you give it a project to do something overnight and metaphorically walk away from it, check back with what it has done in 6 or 8 hours. 0.05 to 0.1 tok/s on the other hand, as reported in the URL for t…
With only 1 PCIe 5.0 SSD, the reading throughput is still significantly more than 10 times faster than on author's system.
So it is likely that inference speeds around 1 token/s are achievable on something like a NUC mini-PC.
Re: Show HN: Getting GLM 5.2 running on my slow computer
#198Re: Show HN: Getting GLM 5.2 running on my slow computer
#199So, could larger models work this way too?
Re: Show HN: Getting GLM 5.2 running on my slow computer
#200Earlier quoted context omitted.
Xeon Scalable in general seems like a good idea due to 6-channel (relatively) inexpensive RDIMM memory, but I've been reading that NUMA kills inference performance. Anyone got experience with multi-socket systems? IIRC even within the socket these cpus are divided into sub-numa nodes.
Even though LLM benchmarks are very opinionated, I would really like to see some numbers for the setup parent suggested. From what I read elsewhere, anything below $40K in HW costs is not worth the effort for coding models locally.
So for optimal speed the models must be quantized in this format.
It is very likely that with INT8 models those CPUs are fast enough so that the inference throughput is limited by the memory bandwidth (384-bit interface to DDR4-2933 per socket, i.e. 282 GB/s for both sockets).
The memory throughput for such an old server is very similar to an AMD Ryzen Strix Halo, NVIDIA DGX Spark or Apple M5 Pro, but it has much more memory.
The inference speed should be very similar to those, but with bigger LLMs.