Earlier quoted context omitted.
plus a special sauce for counting the number of specific bp repeats, due to in-del events, this is not something I am not too familiar, but presumably the number of a specific k-mer repeats you have in these genes of interest might correlate to a specific type of cancer? (would love to hear someone who is an expert in this field their opinion). "Copy number variant" refers to larger deletions and duplications that ca…
Thanks for your detailed explanation. Just out of curiosity and to follow-up, presumably this is a example of the list of detected CNVs in a TCGA Breast Cancer data-set you're referring to: http://cancer.sanger.ac.uk/cosmic/gene/analysis?ln=BRCA1#cnv... According to Sanger (or maybe TCGA?), a gain is when a genomic region (for a diploid) has more than five absolute copies of this region and a loss is when the genomic…
The figure you linked is a good explanation. The split read method is helpful for finding the edges of the CNV, while the number of reads (relative to other regions that were tested) can give an idea of the number of copies. The problem is that these methods all have their own unique biases/noise that makes it non-trivial to figure out the absolute copy number change.
Ideally they would find a similar CNV that has some clinical association.
The DGV has a lot of reference CNVs. Here are some in BRCA1: http://dgv.tcag.ca/gb2/gbrowse/dgv2_hg19/?name=id:3087443;db...