Live data from Hacker News

Researchers generate complete human X chromosome sequence

genome.gov

41–50 of 50 posts

Re: Researchers generate complete human X chromosome sequence

#41

The article doesn’t mention but now I’m curious what kind of read length and error rate they are achieving. This could have huge impacts across all sequencing.

It looks like they're using Oxford nanopore and PacBio sequencing technologies for the long reads. These are two up-and-coming sequencing technologies focused on extremely long reads. My understanding of both is that their error rates on individual base pairs are too high to reliably determine the actual sequence on their own (something like 15% error rates). Typically the long reads from these technologies are used…

15% is quite outdated. There have been major updates to nanopores and software. Typical single read error rate is less than 5% these days.

Single read accuracy is not as important for such projects. As coverage gets to 50-60X, expected assembly accuracy is Q30 on human.

Re: Researchers generate complete human X chromosome sequence

#42
post #17

Earlier quoted context omitted.

From the article: "Repetitive DNA sequences are common throughout the genome and have always posed a challenge for sequencing because most technologies produce relatively short "reads" of the sequence, which then have to be pieced together like a jigsaw puzzle to assemble the genome. Repetitive sequences yield lots of short reads that look almost identical, like a large expanse of blue sky in a puzzle, with no clues…

So 20 years ago when we "sequenced the human genome", we actually didn't? If you'd asked me whether or not we had a complete sequence of an X chromosome before I saw this I would have said, "Of course we have one, for over 20 years".

So much about the original announcements was overhyped PR. The original assembly was super-crappy and super-gappy. The folks running the two projects were exhausted and declared victory, then moved on.

Of all the fields that I've worked in, genomics has been one of the most overhyped (virtual drug discovery is the other) and it takes a ton of training just to understand how messed up the field is.

Re: Researchers generate complete human X chromosome sequence

#43
post #40

Earlier quoted context omitted.

That's right -- sequenced genomes are typically assemblies of short fragments. The assembly algorithms fail in areas of low complexity or when large sequences are repeated an unknown number of times.

So what DID we achieve 20 years ago? Like, what happened then that led to the claim of "human genome has been sequenced" and what is the additional progress that was made now?

What was announced 20 years ago was an incomplete assembly that met certain metrics that made sense at the time. https://www.nature.com/articles/35057062 is the paper. It describes the assembly as a "partial draft".

The section "Background to the Human Genome Project" gives some color on why they did what they did (TL;DR there was an ostensibly competitive race between the public project and a private one).

I ended up providing some useful resources for helping uncover just how bad genomic assemblies were (at the comptuational level): most genomic assemblies using whole genome shotgun sequencing used a number of heuristics which were believed to be correct, but I suspected that the heuristics failed to deal with repetitive regions and short sequences well. So I built a computing system with >1M xeon cores (Google Exacycle) and we provided the system to Gene Myers (who did the original WGS assembly for Celera). he used the system to do an all-vs-all comparison of sequence pairs, which found numerous bugs and problems with the heuristics that were being used. It was a huge amount of compute but the result was that myers was able to use PacBio data to assembly a significantly better genome, faster, on a laptop: (https://www.yuzuki.org/favorite-talk-agbt-2014-gene-myers-ma...)

Re: Researchers generate complete human X chromosome sequence

#44

Dumb question: is there a x_chromosome.txt with the sequence in order? Why do geneticists not talk about it this way?

At a certain level of abstraction, we can treat it that way and it is good enough for many use cases. In biological and physical reality, no.

Each human started with between 1 and 5 copies of the X chromosome. Those copies are different in various ways. Many of the differences are single nucleotide variation, identical in a region but with a single letter changed. There are also tandem repeats where there might be a CAG sequence that occurs one or dozens of times. (Counting the number of repeats like this is often used for DNA fingerprinting.) There is also ample larger-scale structural variation, which includes whole regions of the genome present present or absent in one copy or another, or maybe copied multiple times in a row, or moved in from another chromosome, or reversed.

Complicated enough? On top of that you have to add the fact that there are trillions of cells in each human and in those trillions of cells you will have many slightly different copies of the original 1 to 5 X chromosomes from when that human was a single-cell organism. You will definitely have changes at the ends of the chromosomes, the telomeres, as they are made up of variable tandem repeats. You'll also have single nucleotide mutations, and if you're unlucky, bigger changes. On some chromosomes (not chromosome X), there's also V(D)J recombination, where our immune "memory" is actually encoded in changes to genome sequence in particular cells. Cancer or a pre-cancerous syndrome will increase the frequency and severity of these changes.

If you want to sequence a whole chromosome you have to contend with the fact that the most accurate methods for sequencing generally give you reads of 1000 nucleotides or less each and you have to assemble them together. People liken the problem to putting together a jigsaw puzzle, but it's not like assembling a jigsaw puzzle from a single box. It's more like taking hundreds of boxes of supposedly the same jigsaw puzzle (but in reality some small changes that make things fit together not quite right), dumping them all in a pile, randomly removing a bunch of them, and then trying to figure out how everything fits together. Also there are many parts of this puzzle with identical artwork and that fit together identically! Good luck!

Scientists have been applying a lot of ingenuity to this puzzle for decades and getting a whole chromosome assembly like this is a big milestone.

Re: Researchers generate complete human X chromosome sequence

#45
post #40

Earlier quoted context omitted.

That's right -- sequenced genomes are typically assemblies of short fragments. The assembly algorithms fail in areas of low complexity or when large sequences are repeated an unknown number of times.

So what DID we achieve 20 years ago? Like, what happened then that led to the claim of "human genome has been sequenced" and what is the additional progress that was made now?

For a good read, I recommend The Gene: An Intimate History by Siddhartha Mukherjee

Re: Researchers generate complete human X chromosome sequence

#46
post #42
post #17

Earlier quoted context omitted.

So 20 years ago when we "sequenced the human genome", we actually didn't? If you'd asked me whether or not we had a complete sequence of an X chromosome before I saw this I would have said, "Of course we have one, for over 20 years".

So much about the original announcements was overhyped PR. The original assembly was super-crappy and super-gappy. The folks running the two projects were exhausted and declared victory, then moved on. Of all the fields that I've worked in, genomics has been one of the most overhyped (virtual drug discovery is the other) and it takes a ton of training just to understand how messed up the field is.

Wow, this is news to me. What a farce haha. Thanks for adding clarity here and breaking a bad assumption I had!

Re: Researchers generate complete human X chromosome sequence

#47

Earlier quoted context omitted.

I dont think you can call pacbio up and coming at this point, but nanopore certainly. And those error rate examples are way way too high - illumina is closer to Q30, which is a 1/1000 error rate[0]. 15% would result in an unusable sequence. https://emea.illumina.com/science/technology/next-generation...

The sequencer may report a quality score of 30, but that doesn't guarantee that the error rate when you align to the genome will actually be 1/1000. Still, you're right that good quality Illumina data can do significantly better than 1% error rate. You can't always get "good quality" data, but I imagine that the researchers on this project probably could, given the well-controlled experimental setup. And yes, a 15% e…

Yes the higher error ones are used for alignment- and even then too high an error rate in a very repetitive region (especially depending on the error type - misreads vs skipped bases etc) make it too challenging to build a scaffold to align your illumina reads.

As of 2018 error rate for alignments with nanopore was around 3-6 percent

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6053456/

Re: Researchers generate complete human X chromosome sequence

#48
post #43
post #40

Earlier quoted context omitted.

So what DID we achieve 20 years ago? Like, what happened then that led to the claim of "human genome has been sequenced" and what is the additional progress that was made now?

What was announced 20 years ago was an incomplete assembly that met certain metrics that made sense at the time. https://www.nature.com/articles/35057062 is the paper. It describes the assembly as a "partial draft". The section "Background to the Human Genome Project" gives some color on why they did what they did (TL;DR there was an ostensibly competitive race between the public project and a private one). I ended u…

Ultimately, is there anything you could do with repetitive regions? Like, were the problems with the heuristics or just a mismatch between ~100 base reads and multi-kb repeats?

Re: Researchers generate complete human X chromosome sequence

#49
post #13

Earlier quoted context omitted.

From the article: "Repetitive DNA sequences are common throughout the genome and have always posed a challenge for sequencing because most technologies produce relatively short "reads" of the sequence, which then have to be pieced together like a jigsaw puzzle to assemble the genome. Repetitive sequences yield lots of short reads that look almost identical, like a large expanse of blue sky in a puzzle, with no clues…

> Repetitive sequences yield lots of short reads that look almost identical, like a large expanse of blue sky in a puzzle, with no clues to how the pieces fit together or how many repeats there are Ah, https://en.wikipedia.org/wiki/Clock_recovery ! Too bad DNA isn't a run-length limited code. (Wouldn't that be something.)

DNA is a code, but it's not just a code. It is also a really long molecule that has to bend and fold up in certain ways.

There are error detecting codes, in a way. Protein is encoded by 3 base codes, and if you insert or delete bases not in a multiple of 3 it will be misaligned, then eventually likely encode a stop code and cause the bad protein to be truncated and likely removed via nonsense mediated decay.

Re: Researchers generate complete human X chromosome sequence

#50
post #43

Earlier quoted context omitted.

What was announced 20 years ago was an incomplete assembly that met certain metrics that made sense at the time. https://www.nature.com/articles/35057062 is the paper. It describes the assembly as a "partial draft". The section "Background to the Human Genome Project" gives some color on why they did what they did (TL;DR there was an ostensibly competitive race between the public project and a private one). I ended u…

Ultimately, is there anything you could do with repetitive regions? Like, were the problems with the heuristics or just a mismatch between ~100 base reads and multi-kb repeats?

most people have moved to techniques that produce longer reads with higher error rates. good coverage + longer reads overcomes the higher error rate. You can read the DALIGN paper by Myers to learn more. Or read https://dazzlerblog.wordpress.com/author/thegenemyers/
Post reply on HN