Earlier quoted context omitted.
Surprised no one has commented this but the latency requires the model to be trained in tiny fragments on each device which is currently a field of research that is being explored. As it stands now basically all of a model needs to be loaded into memory. There’s a whole field here and people exploring this problem, colloquially solving this would enable Federated Learning and whoever figures this out will far eclipse…
Where can I read more about this field of research?
https://en.wikipedia.org/wiki/MLOps
Armed with that term, we get (haven’t read):
Machine Learning Operations (MLOps): Overview, Definition, and Architecture
https://arxiv.org/abs/2205.02302
[+ps]
Better resource: https://ml-ops.org/
Based on my (limited) exposure to date, there is tremendous opportunity for software engineers and architects to make impact in ML systems. There is a pronounced lack of seasoned engineering talent (outside of big players like DataBricks, et al) and this knowledge gap sits behind an experience curve that mere IQ can’t jump over. Our experience as software architects and engineers is very valuable.
Know this and recognize the value you will bring to the table.