Live data from Hacker News

The complete sequence of a human Y chromosome

nature.com

191–200 of 232 posts

Re: The complete sequence of a human Y chromosome

#192

Here is the "dumb question" I've always had about recording the human genome. We all have different DNA. So is "the human genome" some kind of "average" DNA, or is it the DNA of whoever they sampled, or is it maybe an overview of what is common for all of us?

In this case it's one haplotype of one cell line.

If you want a population model of genomes you need a pangenome.

See "pangenome graph", "variation graph", and the human pangenome project.

Re: The complete sequence of a human Y chromosome

#193
post #93
post #63

Earlier quoted context omitted.

Lincoln Stein is great, and perl was certainly critical to many processes, but it was and remains fairly niche in genomics, which used much more C++, Java, and later Python. IMHO the person who "saved the genome project" was WJ Kent, who developed the assembler, BLAT, that the public project needed. I strive to point out that he wasn't a sole hero, nor was Lincoln. What I really like about BLAT is that while Celera w…

I had the privilege of working a co-op term at Stein's lab at the OICR. I encountered quite a bit of Perl during my short time there, and have yet to see it elsewhere in the modern enterprise tech world. BioPerl in particular stands out as a fairly substantial project in the bioinformatics space.

Amazon still has some Perl kicking around, mainly on the main website. I don't think there's any new development though, just maintenance and gradual replacement.

Re: The complete sequence of a human Y chromosome

#194

Earlier quoted context omitted.

Some of the research about being able to make simple animals grow structures from other animals in their evolutionary “tree” by changing chemical signaling—among other wild things like finding that memories may be stored outside the brain, at least in some animals—makes me think you need more than just the “code” to get the animal that would have been produced if that “code” were in its full context (of a reproductiv…

My favorite trivia here is that flamingos aren't actually "genetically" pink but "environmentally" pink because they pick up the color from eating algae. Except of course "genetics" and "environment" aren't actually separate things; sure, people's skin color isn't usually affected by their food, but only because most people don't eat colloidal silver. https://en.wikipedia.org/wiki/Paul_Karason

AFAIK most poisonous frogs also aren’t “naturally” poisonous—they get it from diet. Ones raised in captivity aren’t poisonous unless you go out of your way to feed them the things they need to become poisonous.

Re: The complete sequence of a human Y chromosome

#195

Here is the "dumb question" I've always had about recording the human genome. We all have different DNA. So is "the human genome" some kind of "average" DNA, or is it the DNA of whoever they sampled, or is it maybe an overview of what is common for all of us?

They are talking about one 'reference genome'. The variation from human-to-human is relatively small (a few million bases out of 3 billion). The reference genome has historically been some kind of average/mosaic of several individuals (this has obvious disadvantages), good enough to put reads in the right place (mostly), and call 'variants' - the differences that make the test genome unique.

The latest/greatest end-to-end T2T reference is entirely based on 'HG002' an individual from Utah, due partly to new information derived from long read technologies.

Re: The complete sequence of a human Y chromosome

#196
post #104
post #54

Earlier quoted context omitted.

Sure, although I'm not aware of anybody who is contemplating quite the level I believe is necessary to really nail the problem into the ground. When I worked at Google, I proposed that Google build a datacenter-sized sequencing center in Iowa or Nebraska near its data centers, buy thousands of sequencers, and run industrial-scale sequencing, push the data straight to the cloud over fat fiber, followed by machine lear…

But why Google? This is what big pharma are doing. Also you can outsource the data collection part. See for example UK Biobank. Their data are available to multiple companies after some period so it makes it more cost efficient.

Why Google? Because this is a big data problem and Google mastered big data and ML on big data a long time ago. Most big pharma hasn't completely internalized the mindset required to do truly large-scale data analysis.

Re: The complete sequence of a human Y chromosome

#197

Earlier quoted context omitted.

Very similar. The difference between you and a chimp is only 4%, the largest difference between two individual humans is about a tenth of that. The person they picked is pseudonymously HG002, who is an Ashkenazi man who took part in the project and consented to commercial distribution of his genome.

Interesting they chose Ashkenazi, given this ancestry is fairly unique. https://gnomad.broadinstitute.org/news/images/2018/10/gnomad...

Geneticists like insular populations to reduce signal to noise ratio. Another favorite population is a utah population of mormons for similar reasons. African genetics for example are trickier to parse out cause and effect as there is a lot more genetic diversity (and therefore noise that makes it difficult to identify true signal), but they too are studies sometimes for this reason specifically. The Yoruba are pretty well sequenced. Among europeans, a good resource is the icelandic genome project, not only due to the sample size but the nature of the bottlenecked population.

Re: The complete sequence of a human Y chromosome

#198

So I have my whole genome sequenced by Nebula. What do I have to do to match it up?

Can someone explain how "I have my whole genome sequenced by Nebula" relates to the news just now that "The human Y chromosome has been completely sequenced"? How can someone have their whole (!) genome sequenced already when so far we weren't able to fully sequence the Y chromosome. And this person seems to have a Y chromosome.

because commercial "whole" genome sequences aren't really whole. But, they normally deliver the raw reads to you in a 50+GB file so I suppose you could take the reads in that file that don't map to the previous reference and try to map them to the new one. Unless you are an expert it's unlikely you would get any actionable results.

Re: The complete sequence of a human Y chromosome

#199
post #5

For those who don't recall: Back in the Dark Ages, there was a race to decode the human genome. The leading competitiors (wealthiest) were Celera Genomics and the Human Genome Project. After some time, Celera (headed by Craig Venter), announced they had done the deed. However, what Celera had actually done was used what they called a "shotgun method", which meant they took small samples here and there, then built a m…

This comment is needlessly negative. What they did had value and they did not hide what they did either, at least that wasn't my impression. The statistical "shenanigans" is better termed as an innovation.

Re: The complete sequence of a human Y chromosome

#200
post #47

Earlier quoted context omitted.

This is a complex question. The cocktail soup in a gamete (sperm or egg) and the resulting zygote contains an awful lot of stuff that would be extremely hard to replace. I could imagine that if the receiving civilization was sufficiently advanced and had a model of what those cells contained (beyond the genomic information) they could build some sort of artificial cell that could bootstrap the genome to the point of…

I’m just pondering this, and it’s not clear to me that there is anything intrinsic in the genome itself that explicitly’says’ “this sequence of DNA bases encodes a protein” or even “these three base-pairs equate to this amino acid”. I wonder if that information could ever really be untangled by a civilisation starting entirely from scratch without access to a cell

Yes, it's intrinsic in the genome but implemented through such a complicated mechanism that attempting to understand these things from first principles is impractical, not impossible.

In genomic science we nearly always use more cheaply available information rather than attempt to solve the hard problem directly. For example, for decades, a lot of sequencing only focused on the transcribed parts of the genome (which typically encode for protein), letting biology do the work for determining which parts are protein.

If you look at the process biophysically, you will see there are actual proteins that bind to the regions just before a protein, because the DNA sequences there match some pattern the protein recognizes. If you move that signal in front of a non-coding region, the apparatus will happily transcribe and even attempt to translate the non-coding region, making a garbage protein.

Post reply on HN