As an expert in machine learning, I don't see how someone would expect this to work. Actually to predict a disease, you need true positives and true negatives. If you only had access to the true positive data, it would be a lot harder to predict accurately.
1000 patients come in with symptoms that look like cancer at year 0.
100 actually get diagnosed with cancer at some point between year 0 and year 5.
Presumably, the remaining 900 didn't have cancer at year 0.