Pool spare GPU capacity to run LLMs at larger scale
1–4 of 4 posts
Re: Pool spare GPU capacity to run LLMs at larger scale
#2You lost me on "spare GPU". I don't have any capable GPUs, let alone spare ones :)
Re: Pool spare GPU capacity to run LLMs at larger scale
#3This is very promising, definitely looks more user friendly than exo. Can't wait to try it out.
Re: Pool spare GPU capacity to run LLMs at larger scale
#4> MoE models via expert sharding with zero cross-node inference traffic
This makes the whole project questionable