One part that people from the software side tend to underestimate is how fuzzy and analog everything in biology is. Genomics look more predictable and organized at first, but even these parts are quite fuzzy and subject to all kinds of physical effects. I'd strongly recommend in reading up on the parts of cell biology that come after this. Otherwise you'll get the wrong impression of how messy biology actually is.
Introduction to Genomics for Engineers
11–20 of 48 posts
Re: Introduction to Genomics for Engineers
#12This guide is also made from me (or some of the me from a couple years back). I haven't read the whole thing yet and it's probably clearly stated at some point (though one can deduce it with the beginning already) but the surprise for me was that this field is highly statistical. Before starting I had the (very) naive view that it was possible to read the genome as one reads a file and look at what's going on. But th…
But they produce short reads, and because DNA is full of repetitive fragments, it's not always clear where the read came from.
We also have two copies of genes, which also further complicates matters.
The first startup where I worked, developed synthetic long reads on top of Illumina's hardware. We could stitch together 50kbp reads, which really helped with de-novo sequencing.
Re: Introduction to Genomics for Engineers
#13One part that people from the software side tend to underestimate is how fuzzy and analog everything in biology is. Genomics look more predictable and organized at first, but even these parts are quite fuzzy and subject to all kinds of physical effects. I'd strongly recommend in reading up on the parts of cell biology that come after this. Otherwise you'll get the wrong impression of how messy biology actually is.
And most of genomics is still stuck in 1930ties. Many people believe gender is somehow related to genes, which is objectively not true! Or that genes are somehow related to your religion!
There's the X and Y chromosomes, those produce a binary result (unless you have a genetic anomaly). And after that comes the messy and fuzzy parts I mentioned, where those genes trigger changes in hormone levels and development. And those parts are analog, very complex and contain a lot of different parts. So the outcome is not binary anymore.
Re: Introduction to Genomics for Engineers
#14This is a weird description, because ... it is not really "broken up". Each chromosome could be shuffled and put into different cells in different numbers. Now, it is unlikely that the resulting cell would be viable or useful, but my contention here is the "broken up" part. Chromosomes are just a way to handle the genome set. There are reasons why bacteria do not have chromosomes and this has mostly to do with the amount of DNA. To call this "breaking up" is a very strange description. (Size is not the only reason; duplication of the DNA before cell division is another important factor; bacteria usually have just one origin of replication, eukaryotes have several on each chromosome, otherwise the S-phase in the cell cycle would simply take too long.)
> Each genome is a biochemical database that, if properly accessed, can inform how our bodies function.
This is also a very strange description, aka "biochemical database". Not everything in a genome has a role with regards to biochemistry or metabolism. Some is just regulatory RNA; some of this relates to metabolism, but you also have e. g. piwiRNA or silencers of transposons and so forth. That in itself has only very rarely a biochemical function, with some exceptions (e. g. I would classify tRNA as related to metabolism, and many viruses have tRNA or use tRNA as quick-starters, but most of those regulatory RNAs do not have any function for metabolism directly, other than e. g. repurposing energy towards their own reproduction).
To me it seems as if the article was written by an engineer. That's fine, but it also means that the thinking is quite biased. Genetics is not quite so easy to engineer; a good example are leaky promoters used in synthetic biology (just ask the people who use such promoters how to make them un-leaky) or off-target cleavage effects in CRISPR-Cas(9 or whatever is used); I am pretty certain they'll give excuses as to why 100% accurate gene therapy isn't yet ready for the masses. And they'll do that for quite some years to come, I bet, usually hiding behind "it will cost too much" - when in reality, it should cost very little, if it were to work, rather than this just becoming the new meta-milking scheme.
Re: Introduction to Genomics for Engineers
#15It's definitely possible to learn enough to be productive within a few months, but to actually comprehend and understand the underlying biology takes much, much longer. I still don't understand much of what is presented by people from other labs outside of my specialty.
Re: Introduction to Genomics for Engineers
#16This guide is also made from me (or some of the me from a couple years back). I haven't read the whole thing yet and it's probably clearly stated at some point (though one can deduce it with the beginning already) but the surprise for me was that this field is highly statistical. Before starting I had the (very) naive view that it was possible to read the genome as one reads a file and look at what's going on. But th…
For example, sequencing instruments include base quality strings in the output. Base qualities are estimates how likely the instrument got each sequenced base right. But most people don't want to store that much noise, especially when the actual data is highly compressible. So the base qualities get quantized using more or less principled methods that seem to work well empirically.
Read aligners make similar estimates of how likely they got the correct alignment for each read. Those estimates are typically based on simplistic models and a number of assumptions. There are two main components in the estimate. One is based on comparing the primary alignment the aligner chose to the secondary alignments it also found. Another is an estimate that the aligner didn't find the correct alignment, because that part of the sequenced genome is too different from the reference. The latter is obviously handwavy. And the aligner cheats in the former. Because people don't want to wait 10x or 100x longer for better results, the aligner gives up early and estimates how good secondary alignments it might have found if it had actually done the work.
And then there is variant calling. At some point, the state-of-the-art callers were statistical. But then people got better results with neural networks. Or at least the results were empirically better.
Re: Introduction to Genomics for Engineers
#17Earlier quoted context omitted.
And most of genomics is still stuck in 1930ties. Many people believe gender is somehow related to genes, which is objectively not true! Or that genes are somehow related to your religion!
Scientists tend to understand that part, that's more of a political/cultural thing (I'm ignoring the language part here entirely about which terms to use for which concepts here). There's the X and Y chromosomes, those produce a binary result (unless you have a genetic anomaly). And after that comes the messy and fuzzy parts I mentioned, where those genes trigger changes in hormone levels and development. And those p…
Re: Introduction to Genomics for Engineers
#18Re: Introduction to Genomics for Engineers
#19Re: Introduction to Genomics for Engineers
#20> In plants and animals, DNA is broken up into a number of large sequences called chromosomes that are tucked into the nucleus. This is a weird description, because ... it is not really "broken up". Each chromosome could be shuffled and put into different cells in different numbers. Now, it is unlikely that the resulting cell would be viable or useful, but my contention here is the "broken up" part. Chromosomes are j…
The DNA of a nucleated (a.k.a. eukaryote) cell is indeed split into multiple chromosomes and the number of the chromosomes and the number of genes on each chromosomes and the sequence of the genes on each chromosome are normally constant for a species and very similar for closely related species. The cells that have an incorrect number of chromosomes (which happens when a cell division does not work correctly) will normally die soon, because their DNA is incomplete.
Only when a cell has all the normal chromosomes, but also some extra chromosomes, it has a chance to survive, even if the extra chromosomes may interfere with some internal processes. Because superfluous chromosomes are much less harmful than missing chromosomes, there have been cases, rare at animals, but frequent at plants, when the entire genome has been doubled, when some cell division has failed, but the descendants of that cell have survived.
There are animals for which the number of the chromosomes and the sequence of the genes on them has remained unchanged for many hundreds of millions of years, though there are also animals where the DNA has been completely rearranged, because some chromosomes have fused into a single chromosome, other chromosomes have split into multiple chromosomes, and on some chromosomes the genes have been shuffled.
Nowadays, it is known that it is very likely that the common ancestor of all animals except comb jellies had 29 chromosome pairs (humans have 23 pairs). In most animal branches there were more chromosome fusions than chromosome splittings, so in the present most animals have fewer than 29 chromosome pairs.
Among the animals with the most conservative genomes are some sponges, some jellyfish, many echinoderms and some of their relatives, i.e. acorn worms, the lancelets, some nemertean worms and some bivalves.