Incredible how much benefit alphafold has brought. And all of that from a less than 100million parameters model. I might be dumb but could they scale it up and make an alphafold 3 with maybe like 10bln params? Would it be a lot better assuming the same training effort is put into it? If it does, can't biotech companies just go nuts and make a 100bln params internal model and have all the protein structures they want?
Has there been much research into the idea of distributing these large models across many heterogeneous machines? I'm wondering if there could be a path towards a mix of alphafold and folding@home, with donated idle compute resources being used to train/run the models. Designing for that sort of fragmentation could also make it easier to slowly run oversized models on local machines with swapped memory.
Perhaps if there was a way to combine mining and folding, to allow participants to somehow gain a share of the output? Eg each folded protein would have a unique hash, which could then be traded?
And yes, I hate everything about what I just typed.