Live data from Hacker News

Researchers generate complete human X chromosome sequence

genome.gov

21–30 of 50 posts

Re: Researchers generate complete human X chromosome sequence

#21

The article doesn’t mention but now I’m curious what kind of read length and error rate they are achieving. This could have huge impacts across all sequencing.

It looks like they're using Oxford nanopore and PacBio sequencing technologies for the long reads. These are two up-and-coming sequencing technologies focused on extremely long reads. My understanding of both is that their error rates on individual base pairs are too high to reliably determine the actual sequence on their own (something like 15% error rates). Typically the long reads from these technologies are used…

I dont think you can call pacbio up and coming at this point, but nanopore certainly.

And those error rate examples are way way too high - illumina is closer to Q30, which is a 1/1000 error rate[0]. 15% would result in an unusable sequence.

https://emea.illumina.com/science/technology/next-generation...

Re: Researchers generate complete human X chromosome sequence

#22

The article doesn’t mention but now I’m curious what kind of read length and error rate they are achieving. This could have huge impacts across all sequencing.

The "ultra long" nanopore reads used in this study are often greater than 100kbp in length and occasionally up to 1Mbp

Re: Researchers generate complete human X chromosome sequence

#23

In a simple English, could someone explain why is this good and what does it all mean?

Imagine you have a bunch of aerial photographs that you're trying to assemble into one big mosaic of the entire region by matching up the overlaps on the edges. The problem is, this is a desert, and significant portions of the region are just flat expanses of empty sand, so all the aerial photographs from those regions look pretty much identical. Even worse, there's several of these flat sandy regions in the area, so you don't even know which region those photos came from.

So despite having taken multiple photos of every square inch of land in your target area, there's no way you can assemble them into one big image just by matching up the overlaps. Without a source of larger-scale information about the region, like a satellite photograph or GPS coordinates for the photos, you have no way of knowing how wide that desert is. All you know is that it's wider than one or two photographs.

This is essentially same problem that current genome assemblies have: there are regions of repetitive sequence in the genome, so all the sequencing reads from those regions look identical to each other, just like the photographs of flat sandy desert, and there's no way to tell how they're supposed to overlap to form the full sequence. The only way to resolve these regions is with a technology that can read all the way through from one end to the other without stopping, producing a single contiguous sequence.

The link here describes the fruits of an effort using exactly those sorts of long-read technologies to fill in all the gaps in the X chromosome sequence, thus generating a single contiguous sequence from end to end, something that hasn't previously been possible for DNA sequences of this size.

As to why this is important, these repetitive sequences, despite being apparently featureless, still sometimes have important effects (not unlike the apparently dead and featureless desert in the analogy). In addition, sometimes there are "oases" of functionally important non-repetitive DNA sequence within the "desert" of repetition, and previous genome assembly methods would not be able to tell where these oases belonged. All of this is important because many functional DNA elements are cis-acting. That is, they exert effects on genes that are nearby on the genome. So if you don't know where they belong, then you don't know what they're doing.

If you can assemble one big chromosome sequence from end to end, all of the above problems go away, and you can finally get on with the analysis you wanted to do anyway and stop worrying about not being able to calculate meaningful distances between DNA elements.

Re: Researchers generate complete human X chromosome sequence

#24
post #6

So how do you actually isolate one chromosome to sequence it?

Step 1: the researchers use a diploid cell line where the entire diploid X chromosome is homozygous: both copies of the X chromosome in this cell line are identical. This is part of why the researchers chose to look at the X chromosome. > To circumvent the complexity of assembling both haplotypes of a diploid genome, we selected the effectively haploid CHM13hTERT cell line for sequencing (abbr. CHM13) Incidentally, t…

One minor correction - in step 2 the DNA is not amplified as this would reduce the fragment length and also lose the methylation information

Re: Researchers generate complete human X chromosome sequence

#25

Earlier quoted context omitted.

It looks like they're using Oxford nanopore and PacBio sequencing technologies for the long reads. These are two up-and-coming sequencing technologies focused on extremely long reads. My understanding of both is that their error rates on individual base pairs are too high to reliably determine the actual sequence on their own (something like 15% error rates). Typically the long reads from these technologies are used…

I dont think you can call pacbio up and coming at this point, but nanopore certainly. And those error rate examples are way way too high - illumina is closer to Q30, which is a 1/1000 error rate[0]. 15% would result in an unusable sequence. https://emea.illumina.com/science/technology/next-generation...

The sequencer may report a quality score of 30, but that doesn't guarantee that the error rate when you align to the genome will actually be 1/1000. Still, you're right that good quality Illumina data can do significantly better than 1% error rate. You can't always get "good quality" data, but I imagine that the researchers on this project probably could, given the well-controlled experimental setup.

And yes, a 15% error rate does result in a sequence that is unusable for the purposes of actually knowing the sequence. But a bunch of really long reads with 15% error can still be used to resolve the large-scale structure of a sequence, and then the lower-error-rate Illumina reads can be aligned onto this large-scale scaffold in order to resolve the actual sequence. At least, this is my understanding of how these technologies are typically used together, and given the mention of PacBio, nanopore, and Illumina, that seems to be what was done in this case.

Re: Researchers generate complete human X chromosome sequence

#26

Earlier quoted context omitted.

Step 1: the researchers use a diploid cell line where the entire diploid X chromosome is homozygous: both copies of the X chromosome in this cell line are identical. This is part of why the researchers chose to look at the X chromosome. > To circumvent the complexity of assembling both haplotypes of a diploid genome, we selected the effectively haploid CHM13hTERT cell line for sequencing (abbr. CHM13) Incidentally, t…

Great write up! Thanks! Regarding step 1 how can any human have an entire homozygous X chromosome? Also/rather why not just use a male with one X chromosome?

there are long regions on the Y chromosome that are very similar to the X chromosome, which would make the analysis difficult:

https://en.wikipedia.org/wiki/Pseudoautosomal_region

Re: Researchers generate complete human X chromosome sequence

#27

The article doesn’t mention but now I’m curious what kind of read length and error rate they are achieving. This could have huge impacts across all sequencing.

It looks like they're using Oxford nanopore and PacBio sequencing technologies for the long reads. These are two up-and-coming sequencing technologies focused on extremely long reads. My understanding of both is that their error rates on individual base pairs are too high to reliably determine the actual sequence on their own (something like 15% error rates). Typically the long reads from these technologies are used…

Illumina error rates are <<1% (~0.1%), whereas Nanopore with newer basecalling software is 5-10%. With UMIs you can get a consensus error that's also <<1%. The error profiles are indeed different: Illumina generally creates substitution errors, whereas Nanopore has trouble with "homopolymers" -- counting how many of the same letter occur in a row.

Re: Researchers generate complete human X chromosome sequence

#28
post #17

Earlier quoted context omitted.

From the article: "Repetitive DNA sequences are common throughout the genome and have always posed a challenge for sequencing because most technologies produce relatively short "reads" of the sequence, which then have to be pieced together like a jigsaw puzzle to assemble the genome. Repetitive sequences yield lots of short reads that look almost identical, like a large expanse of blue sky in a puzzle, with no clues…

So 20 years ago when we "sequenced the human genome", we actually didn't? If you'd asked me whether or not we had a complete sequence of an X chromosome before I saw this I would have said, "Of course we have one, for over 20 years".

That's right -- sequenced genomes are typically assemblies of short fragments. The assembly algorithms fail in areas of low complexity or when large sequences are repeated an unknown number of times.

Re: Researchers generate complete human X chromosome sequence

#30

Dumb question: is there a x_chromosome.txt with the sequence in order? Why do geneticists not talk about it this way?

It’s a good question. The answer is no, before this study we didn’t have a gapless “x_chromosome.txt”. We did have 97% of it, but there were parts that were missing here and there. In fact, because the answer is no - which admittedly probably seems wild - this work is very important.

Now, there are much more sophisticated answers, and downstream points to be made about graph genomes instead of a reference, etc (which would also get to your point about why geneticists don’t talk about it this way). But, that’s a broader scope.

Post reply on HN