Live data from Hacker News

Software breakthrough radically boosts the speed of nanopore DNA sequencers

newatlas.com

31–40 of 57 posts

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#31

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

I'm not an expert, but I think the PCR page will help: https://en.wikipedia.org/wiki/Polymerase_chain_reaction

basically for PCR (enzyme based) it's sort of "unzipped" down the middle and the base pairs match a specific way which is how it's duplicated.

The rest I can't answer but I hope this was as fun a read as it was for me the first time. It really is insane to think about.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#32

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

1.

The key to this is understanding that the reads or sequences are long, and the slicing happens at somewhat random locations, such that your reads all overlap with each other (of course in reality it's more complicated than that, due to a bunch of long non-coding sequences and repeating sequences, but this will be mostly true for interesting parts of the genome).

Therefore your mental model of A, B, C is not particularly useful. I would replace it with the following example:

QWERTYUIOP ---[gets chopped up into]----> QWERT ERTY UIOP WE QWER TYU QW

Note that I used the terms "read" and "sequence" interchangeably here because the general notion is the same. But these two words refer to two different things.

2.

Histones are proteins. You can get rid of them simply by adding proteases to the mix. Proteases specifically break down proteins and leave nucleic acid sequences intact.

A commonly used protease is Proteinase K, which you can just order online here: https://www.sigmaaldrich.com/life-science/metabolomics/enzym...

You can also buy kits for specific applications: https://www.sigmaaldrich.com/life-science/molecular-biology/...

Once the histones have been digested, the DNA can be extracted, purified, amplified, etc, at your leisure. A full treatment of this topic would fill an entire textbook, so I will spare you the details.

3.

I'm not sure what this question means, but if you can rephrase I might be able to answer.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#33

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

For 1), the average person here might be better equipped to answer than the average biochemist! Sequences are put back together with de bruijn graphs. Small sequences are compiled into continuous segments based on their overlap. The key piece you might be missing is that there are many many copies of the same genome, randomly fragmented so with luck (or careful experimental design) you’ll have sufficient coverage to complete chromosomes. There’s still lots of tricky regions though.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#34
post #33

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

For 1), the average person here might be better equipped to answer than the average biochemist! Sequences are put back together with de bruijn graphs. Small sequences are compiled into continuous segments based on their overlap. The key piece you might be missing is that there are many many copies of the same genome, randomly fragmented so with luck (or careful experimental design) you’ll have sufficient coverage to…

Biochemist here, this is correct. This falls under "bioinformatics" or "computational biology". Most biochemists are more focused on the "wet lab" part of the job, and less on the in silico fun stuff.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#35
post #31

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

I'm not an expert, but I think the PCR page will help: https://en.wikipedia.org/wiki/Polymerase_chain_reaction basically for PCR (enzyme based) it's sort of "unzipped" down the middle and the base pairs match a specific way which is how it's duplicated. The rest I can't answer but I hope this was as fun a read as it was for me the first time. It really is insane to think about.

Biochemist here! The DNA polymerase reaction can only take place when the gene is unwound from the histone in the first place. Therefore the PCR wiki page will not provide you with the answer... The key is the "DNA purification" prep step where you digest the histones with proteases.

See here: https://en.wikipedia.org/wiki/DNA_extraction

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#36
post #18

If anyone is interested in DNA sequencing, I think the WENGAN paper [0] that came out the other day is pretty interesting too. They are using a hybrid input [1] in combination with their software package to reconstruct higher fidelity reference genomes. The big impact from that paper is really summarized by this: > WENGAN assembly of the haploid CHM13 sample achieved a contig NG50 of 80.64 Mb (NGA50: 59.59 Mb), which…

HiCanu achieved that a year ago. WENGAN is a worse assembler with clearly more misassemblies. It amazes me that this level of paper can be published in Nature Biotech.

They even call it out in the paper, I hadn’t given it a thorough enough reading. Damn, they must’ve been a bit miffed.

I think the novelty they're pitching is the lower resources, but I wonder how much of a problem that is in practice.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#37
post #18

Earlier quoted context omitted.

HiCanu achieved that a year ago. WENGAN is a worse assembler with clearly more misassemblies. It amazes me that this level of paper can be published in Nature Biotech.

They even call it out in the paper, I hadn’t given it a thorough enough reading. Damn, they must’ve been a bit miffed. I think the novelty they're pitching is the lower resources, but I wonder how much of a problem that is in practice.

Several recent assemblers are faster than WENGAN. For nanopore, there are shasta and wtdbg2 (both published). For HiFi, there are Peregrine and hifiasm (both unpublished but with preprints).

I also found their phrasing here misleading: "The run time of WENGAN was at least 183 times faster than that of CANU (UL), while at the same time using less memory than other assemblers such as FLYE and SHASTA". Why not say the other way around (which is also a fact): The run time of WENGAN was at least several times slower than that of shasta, while at the same time using more memory than CANU. Come on, this is not the right way to do comparisons!

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#38

I would be eternally grateful to a bio-chem person to explain two facets of DNA sequencing I've never understood. To motivate the question a short quote from Wikipedia: "For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and a…

If you're interested in how 1 works in details, http://rosalind.info/problems/tree-view/ has a few bioinformatics coding problems. Starting completely from scratch, so you can learn as you go. The reconstruction problem is there as well.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#39
post #24

Earlier quoted context omitted.

to me nanopore seq is not that promising. it has been here for a long time now and the benefits didn't convince a lot of customers to adopt it. there are 2 main use cases now in precision medicine: rare disease and cancer. for both you need high precision reads, which nanopore doesn't provide.

I think you're stuck in the biotech == people mindset. On the bacterial side, it is fantastic for quickly sequcinging and closing genomes. Personally, I think the killer use case is field portable and real time sequencing of pathogens. I've worked with groups (.gov and private, defense and health related) that want to put minION + flongles to use in applications like early detection for bio terrorism and pathogen sur…

I have a colleague who is working on a portable field kit with a minION + laptop + car battery with the intent of being able to sequence and identify pathogens directly in the field, even with no electric grid.

Turns out the hardest part is the sample prep. For bioinformatics, she'll just do a simple kmer mapping against a curated database of pathogen genomes.

Re: Software breakthrough radically boosts the speed of nanopore DNA sequencers

#40
post #14

Earlier quoted context omitted.

to me nanopore seq is not that promising. it has been here for a long time now and the benefits didn't convince a lot of customers to adopt it. there are 2 main use cases now in precision medicine: rare disease and cancer. for both you need high precision reads, which nanopore doesn't provide.

Should the long read lengths allow error correction to work well if there is sufficient coverage?

Only somewhat, because the errors are systematic, and not random. Using the R9 pore flowcells, I've basically given up on getting correct consensus sequences even for influenza genomes (without manual correction, which is too labour intensive). Perhaps the new R10 pore, much better at homopolymers, will solve the problem.
Post reply on HN