Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

471–480 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#471
post #416
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

Then we get the really fun question: if the experimentally determined structure is only 90% accurate, can machine learning actually reach 100%? Can you learn exact truth from inexact examples? Which gets into the concept of whether the ML model has actually learned some deeper conceptual ideas than we have, some deeper truth about how this works. If so, can we somehow extract that truth, or is it truly a black box th…

If you have an experimental error that is somewhat normally distributed around the mean, the the AI should, with enough examples, learn what the rules are that are closest to the mean. Because it will minimize the sum of errors.

So i do think the results could be more accurate than measurement.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#472
post #168

Earlier quoted context omitted.

The full chain is DNA -> mRNA -> Ribosome -> tRNA combinations -> amino acid chain -> protein. It's true that in nature there are many steps between DNA and proteins (this list doesn't even include the steps that mediate the translation, ie. start it, stop it, slow it down, ...), but the structure of a protein is fully determined by the DNA code. Protein folding is about you start from the DNA code that is fed into t…

Thank you very much; almost forgot I did a Phd on the subject ;-) But anyway your answer does not contradict my statement. What you say belongs to the basics of molecular biology, but does not justify that DNA should be considered when determining the structure of proteins. In practice, the amino acid sequence is always already present.

For the sceptics: if you read the referenced article, you will see that it is about protein structure determination by means of deep neural networks. It's not about gene expression, which is a different topic. What benefit does it have to respond to the question "What are the immediate real-world applications of this" (see above) by reciting some molecular biology dogmas from text books mixed with misconceptions, instead of responding to the real question?

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#476
post #459

Earlier quoted context omitted.

We don't know too much about the exact model they made but it looks sufficiently generalizable to be able to give a candidate protein structure for any given sequence. It doesn't automatically cure cancer and inject the drug but that by itself is an amazing tool that if available to everyone will revolutionize biology experimentation. I will definitely blame the protein structure field in multiple levels though. It w…

"We don't know too much about the exact model they made but it looks sufficiently generalizable to be able to give a candidate protein structure for any given sequence. It doesn't automatically cure cancer and inject the drug but that by itself is an amazing tool that if available to everyone will revolutionize biology experimentation." They say on their own press-release page that side-chains are a future research p…

Of course, they havent solved everything, but you seem to be doing exactly what I accuse that entire field (and academia in general) of doing - which is to insist a problem is intractable or hard and undermine someone potentially challenging that. When they released the 2018 results tbey field did embrace it (for sure I'd consider the groups organizing CASP as at least forward thinking) but was still skeptical on how much more progress it can make; now they blow everyone's minds again by a monumental leap and again people want to come say of course this is the last big jump!

I understand the self preservation instincts that kick in when there's a suggestion that the entire field has been in a dark age for a while, but I hope you can see that there might be something fundamentally wrong with how research is done in academia and that is to blame for why this didn't happen sooner, and why it's so hard for many to embrace it.

Regarding your comments on the inapplicability of this current solution for docking, I'm sure that's the next project they're taking up, and let's see where that goes.

This is exactly the same type of progression that happened with Go, where when their software bet a professional player everyone's like "yeah but I bet he wasn't that good". A few years later and Lee Sedol just decided to retire. I am interested to see what happens to that entire academic field in a similar vein, though my interests are more in knowing how science can advance from more people thinking this way.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#478

Fascinating! AlphaFold (and other competitors) seem to use MSA (Multiple Sequence Aligment) and this (brilliant) idea of co-evolving residues to build an initial graph of sections of protein chain that are likely proximal. This seems like a useful trick for predicting existing biological structures (i.e. ones that evolved) from genomic data. I wonder (as very much a non-biologist), do MSA-based approaches also help u…

This is a really insightful question and I need to take some time to fully understand the ensuing discussion.

If my speculation is correct, then drug discovery should use a process of genetic programming, using something like this to score the resulting amino acid sequences. I'm wondering if an artificial process of evolution would be sufficient to satisfy the co-evolution assumption here.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#479
post #53

Pretty interesting that they only used about $15k worth of resources (retail price) to achieve this. It's not a technique that would have been out of reach for other organizations based only on not being able to afford the compute.

I'm pretty sure that this took more than 1 junior engineer-month.
Post reply on HN