Earlier quoted context omitted.
In general it's a (computationally) hard problem to restitch a genome together. Even today, when you 'get your genome sequenced' you are not getting a full read-through of your entire genome's data. Imagine you want to reconstruct the data on two RAIDs that are mostly, but importantly not exactly, mirrors of each other. Each RAID has 23 drives. Each drive has ~1Gb or so of data. And much of the data is not only mirro…
hmm I don`t think so I fathom the complete complexity of the process but with so many powerful GPU`s out there, is there a possibility of reconstruction in a matter of days if not hours?
If you are further interested in parallelization schemes for genomic pipelines, please have a look at our paper on the strengths and limitations of big data technology for genomic analysis (published last week) - https://people.cs.umass.edu/~aroy/sigmod17-roy.pdf