Live data from Hacker News

The complete sequence of a human Y chromosome

nature.com

21–30 of 232 posts

Re: The complete sequence of a human Y chromosome

#22
post #6

Could someone explain exactly what it means to be "completely sequence" the human genome when all humans have distinct genetic makeup (ie, different sequences of nucleobases in their DNA/RNA)?

I've never understood this either. I assume the genome is many megabytes of [ATCG]+. If we have that sequence, what does it tell us? Do we look at it and say "Ah, yes, ...ATGCTACGACTACGACTAGCG... very interesting?"

It’s actually about 3 gigabases (ATCG). There are some recurrent features of the genome whose function we’ve worked out. For example the TATA box is a classic sequence that typically indicates the start of a part of the genome that codes for a protein. The vast majority of the genome doesn’t code for proteins. The function of these genome regions are much more murky. Some of these regions function like scaffolds for proteins to assemble into complexes. These protein complexes then start transcribing the genome into into mRNA. So the genome regulates its own expression, in a sense. Many of the sequences that function in this way are known. There are also just a bunch of parts of the genome that probably don’t do anything. There are also many regions of the genome that are basically self replicating sequences. They code for proteins that are capable of inserting their own genetic sequence back into the genome. These are transposons.

In short, a lot of very painstaking genetics and molecular biology work has gone into characterizing the function of certain sequences.

Re: The complete sequence of a human Y chromosome

#23
post #5

For those who don't recall: Back in the Dark Ages, there was a race to decode the human genome. The leading competitiors (wealthiest) were Celera Genomics and the Human Genome Project. After some time, Celera (headed by Craig Venter), announced they had done the deed. However, what Celera had actually done was used what they called a "shotgun method", which meant they took small samples here and there, then built a m…

[flagged]

Re: The complete sequence of a human Y chromosome

#24

This is actually kind of a huge deal since it means that all 24 chromosomes have now been fully sequenced. As it says in the abstract, up until now the Y chromosome proved difficult to sequence due to its nature.

As others have asked in this thread, what does "fully sequenced" mean to the layman?

It looks like from the preprint that they sequenced the Y chromosome of HG002, which was one of the original 1000 genomes samples from way back in the day, still held in deep freeze at a number of biobanks.

Short-read sequencing data is a notoriously bad datatype for reconstructing the low-complexity / repetitive regions of genomes, so up until recently the most commonly used reference genomes have left many of these regions "dark". According to the preprint, the Y chromosome has the highest density of these low-complexity regions. It's also something of a bioinformatic nuisance when constructing a generic human reference genome, as it's only present in 50% of the population.

Re: The complete sequence of a human Y chromosome

#25

This is actually kind of a huge deal since it means that all 24 chromosomes have now been fully sequenced. As it says in the abstract, up until now the Y chromosome proved difficult to sequence due to its nature.

As others have asked in this thread, what does "fully sequenced" mean to the layman?

It provides a complete baseline/reference DNA "map". Common "problems" show up on specific parts of this map. So you can sequence small subsets of a patient's DNA and compare it to these "problem areas" to detect genetic diseases.

Re: The complete sequence of a human Y chromosome

#27
post #23
post #5

For those who don't recall: Back in the Dark Ages, there was a race to decode the human genome. The leading competitiors (wealthiest) were Celera Genomics and the Human Genome Project. After some time, Celera (headed by Craig Venter), announced they had done the deed. However, what Celera had actually done was used what they called a "shotgun method", which meant they took small samples here and there, then built a m…

[flagged]

I'm sorry, are you saying the media, whose job it is to report and not serve as experts, are useless ignorants, or Celera, which lied to the public, are useless idiots?

If the first, time to let that dead dog lie. It's a tired trope.

Re: The complete sequence of a human Y chromosome

#28

Earlier quoted context omitted.

I've never understood this either. I assume the genome is many megabytes of [ATCG]+. If we have that sequence, what does it tell us? Do we look at it and say "Ah, yes, ...ATGCTACGACTACGACTAGCG... very interesting?"

It's just about 3 gigabytes (each byte a letter). Pretty mind-blowing, if you ask me.

[deleted]

Re: The complete sequence of a human Y chromosome

#29
post #27
post #23

Earlier quoted context omitted.

[flagged]

I'm sorry, are you saying the media, whose job it is to report and not serve as experts, are useless ignorants, or Celera, which lied to the public, are useless idiots? If the first, time to let that dead dog lie. It's a tired trope.

Not a tired trope because that dog isn't dead. Until the media either 1. has no influence or 2. stops being dishonest then it needs to constantly be called out and berated.

Re: The complete sequence of a human Y chromosome

#30
post #5

For those who don't recall: Back in the Dark Ages, there was a race to decode the human genome. The leading competitiors (wealthiest) were Celera Genomics and the Human Genome Project. After some time, Celera (headed by Craig Venter), announced they had done the deed. However, what Celera had actually done was used what they called a "shotgun method", which meant they took small samples here and there, then built a m…

Yes, but the HGP didn't make a full sequence either, using their scaffold-based contig assembly. Both groups declared a truce and announced they were "finished with the first draft"

35 years ago: https://www.nytimes.com/1987/12/13/magazine/the-genome-proje...

33 years ago: https://www.nytimes.com/1990/06/05/science/great-15-year-pro...

29 years ago: the competition gets fierce https://www.nytimes.com/1994/02/22/science/scientist-at-work...

24 years ago: the race is alive! https://www.nytimes.com/1999/03/23/science/who-ll-sequence-h...

23 years ago: draft is complete https://archive.nytimes.com/www.nytimes.com/library/national...

20 years ago: https://www.nytimes.com/2003/04/15/science/once-again-scient...

2 years ago: https://archive.nytimes.com/www.nytimes.com/library/national...

The celera approach using shotgun was tricky because assembling the bits required a great deal of computational finesse and horsepower; "The final assembly computations were run on Compaq’s new AlphaServer GS160 because the algorithms and data required 64 gigabytes of shared memory to run successfully."

At the time the GS160 was the monster machine, a classic big iron UNIX, which could be tightly clustered- unified IO/filesystem between all the machines. The primary author of the shotgun assembly was Gene Myers, who previously had written BLAST, and invented the Suffix Array with Udi Manber.

The public project had its own issues, as many of the teams assigned to work on it still had the "cottage industry/artisinal/academic" approach, then Eric Lander came along and turned it into an industrial process, and parlayed that into running the Broad Institute, a privately funded MIT/Harvard research institute in Boston.

These days petabytes of sequence data are generated every day and stored in clouds. The genome has been a fundamental tool for shaping our studies of humans, although its true potential for understanding complex phenotypes remains elusive.

Post reply on HN