Earlier quoted context omitted.
It isn't really that much of a problem. If you collect data on a large enough cohort of symptomatic individuals, you can recover separate subtypes through cluster analysis.
Cluster analysis will require thousands of not tens of thousands of cases to study. And since we don't know even how many diseases there are, this will have to be an unsupervised CA, so for each case you'll be looking tens if not hundreds of parameters. Machine Learning techniques are well suited to crunching this type of data, but don't underestimate the difficulty in collecting so much data.
I've found that in biological systems, individual variation can be as large as population variation, particularly for something "reactive" like the immune system, which is tied into stress response.