Live data from Hacker News

No human genome has ever been completely sequenced

statnews.com

51–60 of 96 posts

Re: No human genome has ever been completely sequenced

#51
post #49
post #12

Earlier quoted context omitted.

https://youtu.be/fCd6B5HRaZ8 is the best visualization of how the most popular type of DNA sequencer works (that I've found). Imagine you have a string of length 3 billion made by randomly choosing from 4 characters. Like this dna = ''.join(random.choices('atgc', weights=[30.9, 29.4, 19.9, 19.8], k=3_234_830_000)) you get to randomly sample 1 billion[3, page 7] overlapping substrings of length 200[3, page 7] with .1%…

This problem is called sequence assembly or genome assembly, not sequence alignment (which is a related problem).

[deleted]

Re: No human genome has ever been completely sequenced

#52
post #27

Earlier quoted context omitted.

Side note: your comment is one of the few times I've seen the code tag used for actual code on HN, rather than quotations or just indentation.

Yeah, why does Hacker News not have a real way to indent things? It wouldn't make loading the page any less light-weight.

Seriously. I can't go a week in HN without reading someone's complaint about not being able to read a quote, probably on mobile. If users keep complaining regularly, it's a problem with the software interface, not the users.

Re: No human genome has ever been completely sequenced

#53
post #5
post #3

It's strange that this article ends as an advertisement for PacBio sequencing (which can ~50k-60k base reads) but makes no mention of Oxford Nanopore (which has gotten megabase reads and keeps improving). Single molecule nanopore sequencing is on track to sequence across the centromeres of human chromosomes in the next few years.

As someone who's only casually interested in this field, what are the prospects of any of these upstarts actually unseating Illumina from the throne?

The simple explanation is that they are different products that only overlap to a limited degree.

If you're doing a standard RNA-seq experiment, short 100bp to 200bp Illumina reads are just dandy for what you need. Most of the time, at least; you could, for example, be interested in different splice variants for a gene family; in that case longer continuous reads are valuable. But then it's no longer a standard RNA-seq.

PacBio and pore-type sequencing is most valuable for polishing reference genome assemblies. For instance, this is a huge issue for crop plants because they have huge, repetitive genomes and are often polyploid. So there's a lot of interest in creating high quality assemblies not just for e.g. corn, but the major research inbreds (B73, Mo17, etc.) as well as private breeding lines. Some of the most valuable crop QTLs (R genes, microRNAs, transposon fragments, etc.) are just 'junk' sequence that plays a regulatory or defense role, so it's super important.

In general, the Illumina data is very high quality and very inexpensive, and works beautifully provided you have a high-quality assembly to work with, and what you're interested in isn't very repetitive. It has many good applications, so I don't see much risk at this point.

If, at some point, longer reads can be done with high fidelity and with similar cost, then there would be an issue. I don't think that's going to happen anytime soon, but I'm sure Illumina is looking at that long-term.

Re: No human genome has ever been completely sequenced

#54
post #50
post #39

Earlier quoted context omitted.

The problem is Oxford Nanopore data has about a 30% insertion and deletion error rate. While PacBio is less than 1% indel and substitution error. Don't get me wrong, the Minion is amazing, but it can't compare to PacBio in terms of quality... yet.

The error rate of recent nanopore data is much less than 30%. See here for a recent benchmark: https://github.com/rrwick/Basecalling-comparison

Interesting, I haven't tried it on the newest chemistry but last time I did a flowcell (earlier this year) that was the INDEL error rate we achieve.

Substitution error rate was very low, but you should evaluate INDELs specifically, rather than just identity like you did on your git, because it is the major drawback of this technology.

Re: No human genome has ever been completely sequenced

#55
post #43
post #33

Earlier quoted context omitted.

True, and unfortunately it does not render it properly, it‘s cut after 20 chars or so

You should be able to scroll it.

Thanks, you’re right. I wasn’t aware of that. The experience is rather crappy though. Cumbersome to scroll on a smartphone screen and I’d love to see to entire statement not parts of only at a time.

Re: No human genome has ever been completely sequenced

#56
post #42

Earlier quoted context omitted.

The fact is, unfortunately, that Nanopore sequencing (and also, from what I’ve read, PacBio) has a dramatically higher error rate than Illumina sequencing-by-synthesis. In the near future, anyway, I would expect to see inaccurate PacBio/Nanopore long reads being used as scaffolds for accurate Illumina short reads (in fact, this is already happening). Illumina won’t be going anywhere any time soon.

>The fact is, unfortunately, that Nanopore sequencing (and also, from what I’ve read, PacBio) has a dramatically higher error rate than Illumina sequencing-by-synthesis. This is true, but only for Insertions/Deletions. The substitution error rate is comparable to Hi-Seq. So it's good for resequencing but without a reference or decent scaffolds, you're in the dark.

I've only sequenced a few samples on a MinION. The dominant error seemed to be miscounting the length of a homopolymer stretch (e.g. CCC->CCCCC) but the substitution error rate was also clearly worse than Illumina. I've gotten used to less than 0.1% error from DNA sequencing on a HiSeq 2500. My guess is that the MinION data had 1-3% substitution errors.

Re: No human genome has ever been completely sequenced

#57
post #54
post #50

Earlier quoted context omitted.

The error rate of recent nanopore data is much less than 30%. See here for a recent benchmark: https://github.com/rrwick/Basecalling-comparison

Interesting, I haven't tried it on the newest chemistry but last time I did a flowcell (earlier this year) that was the INDEL error rate we achieve. Substitution error rate was very low, but you should evaluate INDELs specifically, rather than just identity like you did on your git, because it is the major drawback of this technology.

Just to be clear this isn't my analysis - and yes most of the errors here are due to indels.

Re: No human genome has ever been completely sequenced

#58

One of the things the HGP revealed is that we have way fewer genes than anyone expected. Given that the amount of data encoded in genes appears to be far too small to account for the sophistication and variability in humans, it strongly implied that large areas of DNA that were previously considered marker/filler DNA were of great epigenetic importance.

It always seemed insanely stupid to me every time I read something about people claiming certain genes being "non-important" or "non-functional"

We have about a 0.0001% comprehension of this whole system, how can you possibly be so arrogant as to think you KNOW that certain parts of the sequence are not important or used in _any_ way? Then again, I'm definitely not qualified to comment on the topic.

Re: No human genome has ever been completely sequenced

#59
post #3

It's strange that this article ends as an advertisement for PacBio sequencing (which can ~50k-60k base reads) but makes no mention of Oxford Nanopore (which has gotten megabase reads and keeps improving). Single molecule nanopore sequencing is on track to sequence across the centromeres of human chromosomes in the next few years.

How many sequencers has Oxford Nanopore sold? It seems like it has perennially been a “in the next few years” technology.

Nanopore sequencing is definitely "already here" in the sense that they're selling a lot of flow cells and people are doing useful research with ONT instruments but it's also a niche technology compared with Illumina.

Re: No human genome has ever been completely sequenced

#60

> * A gene called ARHGAP11B, which was created by one such duplication, causes the cortex to develop the myriad folds that support complex thought; SRGAP2C, also a duplication, triggers brain development.* A question I've never seen addressed: is one justification for junk DNA to create space for beneficial mutations? Obviously some mutations are actively harmful, but the reason most mutations are destructive is that…

Not stupid at all.

If you're interested in that, I'd look up "Whole Genome Duplications". For instance, corn underwent a WGD about 7 million years ago (IIRC, the number might be off), and we can identify an 'a' and 'b' genome. After these events, most of the duplicated genes will be lost, either through purging selection or drift. However, some will stick around and take on new subset functions (see nkrumm reply).

If you really want to get into the weeds on this, there are many metrics that could, in theory, affect the 'evolvability' of a genome. Higher transposon (or any sort of repetitive element) content can lead to large-scale shuffling of genome structure, such as inversions, duplications, etc. The ecological success of the grass family, for instance, may have it's roots in genome structure.

Now, there's not a whole lot we can say about this because it's not something we have a really good grasp of. It's mostly some hypotheses that are quite difficult to test, but this is some of the most interesting genomic research going on, in my opinion.

All of this is quite difficult to understand, however, as it really depends on epigenetic mechanisms like chromatin state, dicers, RISC, and all that. Figure that out, and you'll realize there's no 'junk' DNA, just various non-coding sequence that's part of this larger-scale genome evolution.

Post reply on HN