Earlier quoted context omitted.
What is simplex ethernet?
Imagine two computers A and B. A has two NICs, A1 and A2. B has B1 and B2. So 4 NICs total. You connect directly A1 to B1 and A2 to B2 with crossover cables. You then route all the traffic from A to B over A1 to B1 and all the traffic from B to A over B2 to A2. Why do you do all this? To avoid collisions and the loss of effective bandwidth from back-offs. It only really works with 2 computers because if you add a 3rd…
AMD Strix Halo RDMA Cluster Setup Guide
61–70 of 90 posts
Re: AMD Strix Halo RDMA Cluster Setup Guide
#62Earlier quoted context omitted.
I had a Strix Halo laptop with 128GB which unfortunately died last week. I paid 2800 euro for it. If I buy the same machine today, the sticker price is 7899. The device was not perfect by any means, but the ability to run fairly large models is some kind of magic.
>sticker price is 7899. It's not even worth it at that point. You can get a used enterprise grade SXM baseboard with 4-8 V100/A100 GPUs off eBay at a similar price. That will even get you actual HMB ram and NVlink. Along with 10x the AI performance, assuming you don't care about your electricity bill of course.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#63Hmm, coing PCIe -> NIC -> NIC -> PCIe seems a bit silly, couldn't both devices communicate directly over PCIe?
Re: AMD Strix Halo RDMA Cluster Setup Guide
#64Earlier quoted context omitted.
Ah yea after watching one of the creators youtube videos I realize these benchmarks are combining prefill and decode which isn't super helpful - it seems this struggles with the exact same bottlenecks as all strix halo setups, memory bandwidth. It seems this is still significantly slower than equivalent memory sizing on Mac hardware.
How are the memory bandwidths specs of Macbooks vs this?
Re: AMD Strix Halo RDMA Cluster Setup Guide
#65Earlier quoted context omitted.
>sticker price is 7899. It's not even worth it at that point. You can get a used enterprise grade SXM baseboard with 4-8 V100/A100 GPUs off eBay at a similar price. That will even get you actual HMB ram and NVlink. Along with 10x the AI performance, assuming you don't care about your electricity bill of course.
You can get a new M5 Max MacBook pro with 128 GB unified ram (targeted by Antirez for DwarfStar4) even after the Apple price increases, it's less than 7899 by at least $1000. And you probably won't pull more than 100 Watts.
It isn't a problem for me, more amusing than anything else (I run in Low Power mode 90% of the time) but worth knowing for anyone thats thinking about pushing the hardware to its limit 24/7.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#66I have two 128gb Strix Halos and have been extremely excited about Antirez's (Redis author) work on DS4, especially with 4bit quant using two machines: https://github.com/antirez/ds4 Right now the speed isn't good for GLM 5.2, Deepseek V4 Flash speed is okay for me (actually reading the output) and quite usable. See kyuz0's great recent video here: https://www.youtube.com/watch?v=PkKXm_mKCCM With a bit more speed and…
>The biggest problem is all the tech companies making consumer hardware completely unaffordable, and I don't think this is accidental. Look at Micron's profits and share price lately... You realize "tech companies" isn't a monolith? Micron charging inflated prices doesn't magically benefit OpenAI. The "high prices keep out competitors" theory doesn't make much sense either. It's like saying Dennys benefits from highe…
I think that realistically, companies compete against each other as individuals and compete against smaller companies and individuals acting more like cartels/monopolies, and that's what OP is referring to in terms of hardware purchasing/contacts/pricing. This also extends outside of tech to investing, so it's likely not just tech responsible for this.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#67Earlier quoted context omitted.
>sticker price is 7899. It's not even worth it at that point. You can get a used enterprise grade SXM baseboard with 4-8 V100/A100 GPUs off eBay at a similar price. That will even get you actual HMB ram and NVlink. Along with 10x the AI performance, assuming you don't care about your electricity bill of course.
You can get a new M5 Max MacBook pro with 128 GB unified ram (targeted by Antirez for DwarfStar4) even after the Apple price increases, it's less than 7899 by at least $1000. And you probably won't pull more than 100 Watts.
The cheapest 128GB Macbook Pro here costs €7.949,00.
No doubt a better value than the HP, and will depreciate a lot less quickly, but just as expensive. Unfortunately, not being able to run Linux is a breaking point for me.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#68Earlier quoted context omitted.
I had a Strix Halo laptop with 128GB which unfortunately died last week. I paid 2800 euro for it. If I buy the same machine today, the sticker price is 7899. The device was not perfect by any means, but the ability to run fairly large models is some kind of magic.
>sticker price is 7899. It's not even worth it at that point. You can get a used enterprise grade SXM baseboard with 4-8 V100/A100 GPUs off eBay at a similar price. That will even get you actual HMB ram and NVlink. Along with 10x the AI performance, assuming you don't care about your electricity bill of course.
I didn't get a Strix Halo laptop because it was the best bang per buck, I got it because it was an awesome machine that could do a little bit of everything, fit in a backpack and only needed 140W.
But noone should buy one at 7899, obviously. It was a tough sell for me at the old 2800 pricing.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#69Earlier quoted context omitted.
How are the memory bandwidths specs of Macbooks vs this?
I looked it up: 512 GB/s for the two node AMD cluster, Macbook Pro with M5 CPU has 153 GB/s. But you can get faster Macs with M5 Pro or M5 Max.
Re: AMD Strix Halo RDMA Cluster Setup Guide
#70I have two 128gb Strix Halos and have been extremely excited about Antirez's (Redis author) work on DS4, especially with 4bit quant using two machines: https://github.com/antirez/ds4 Right now the speed isn't good for GLM 5.2, Deepseek V4 Flash speed is okay for me (actually reading the output) and quite usable. See kyuz0's great recent video here: https://www.youtube.com/watch?v=PkKXm_mKCCM With a bit more speed and…