Earlier quoted context omitted.
Distributed, shared memory machines used to do exactly that in HPC space. They were a NUMA alternative. It works if the processing plus high-speed interconnect are collectively faster than the request rate. The 8x setups with NVLink are kind of like that model. You may have meant that nobody has a stack that uses clustering or DSM with low-latency interconnects. If so, then that might be worth developing given prior…
I think existing players will have trouble developing a low latency solution like us whilst they are still running on non-deterministic hardware.
It would great if a company or others with AI hardware were willing to do production runs of chips sold at cost specifically to make open, permissive-licensed models. As in, since you’d lose profit, the cluster owner and users would be legally required to only make permissive models. Maybe at least one in each category (eg text, visual).
Do you think your company or any other hardware supplier would do that? Or someone sell 2500 GPU’s at cost for open models?
(Note to anyone involved in CHIPS Act: please fund a cluster or accelerator specifically for this.)