I am sorry I did not understand all of it. But, would this allow running large MoE LLMs on a local network with experts spread out over multiple cheaper GPUs (or even CPUs)? This would perhaps be more useful than over the Internet, within offices for example.
Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
11–20 of 25 posts
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#12What if one wishes to use various busybox nodes within the house? All the iot devices contributing to matmul but within a LAN?
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#13Cool idea. How do you handle temperature in the verification?
Ho thanks for the comment. Verification does not depend on temperature. Expert execution is deterministic (pure matmul). LUMABRI_VERIFY=N re-runs N% of the calls on a second replica and requires byte-identical output. Temperature (and sampling) happens only on the chatter, after the experts return their activations. So it can be any value (0, 0.7, 1.2…) without affecting the verification contract.
Isn't that only true in theory but wrong in practice due to floating points?
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#14Without diving into an experiment myself, it would be amazing if you could add some stats or experiment logs, if it's not too much problem and you have them.
For example:
Given model XYZ, every assuming 5 donors with a uniform 32GB each, each forward pass shunts xGB over the link. Each pass takes nMS, etc. etc. Resulting in n T/s, assuming latency of n ms.
Do you have such stats? Or perhaps I missed them in the repo?
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#15Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#16Great work. Love the idea. I have been thinking along the same way, but for training. For inference, the waiting time might be turn off for users
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#17Earlier quoted context omitted.
Ho thanks for the comment. Verification does not depend on temperature. Expert execution is deterministic (pure matmul). LUMABRI_VERIFY=N re-runs N% of the calls on a second replica and requires byte-identical output. Temperature (and sampling) happens only on the chatter, after the experts return their activations. So it can be any value (0, 0.7, 1.2…) without affecting the verification contract.
> Expert execution is deterministic (pure matmul). Isn't that only true in theory but wrong in practice due to floating points?
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#18Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#19The biggest benefit I see is to enable RAM constrained GPUs to perform inference of large parameter models with surprisingly high throughput. Because only a single expert is resident, the memory to compute ratio over the network is limited only by the activations, not the weights. For an Moe like kimi k3 where active parameters are 103B, we might expect to achieve performance limited only by ~5 effective tok/s per Tflop and ~1 tok/s per 100GB/s.
The more members of the network, the smaller your resident parameters are required to be. I’m not sure whether we can split layer inference into arbitrary chunks, but if so you’d be able to increase memory throughput by storing everything in GPU caches.
Of course, we expect latency to be relatively high, but that’s a tradeoff that's fine for certain circumstances.
I’m not sure whether there are any issues more with this idea, but it’s a fun one nonetheless :)
Re: Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
#20On the surface it seems similar to https://meshllm.cloud/ , but now I notice that I don't understand either.