Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

111–120 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#111
post #89

Earlier quoted context omitted.

"A few" does appear quite dismissive of the enormous amounts of effort in structural biology so far. There are more than 170,000 structures in the PDB right now. To determine potential targets for drugs we have to understand what the proteins do. Having the structure is not really enough for that, it doesn't tell you the purpose of the protein (though it certainly can give you some hints). In most cases the proteins…

170k is "a few" compared to 180 million (i.e. the size of the PDB as soon as someone runs AlphaFold over everything in the UniProt.) > In most cases the proteins were determined to be interesting by other experiments, and then people decided to try and solve their structure. Yes, that's what we're doing right now , because structure is not a useful predictor, because we don't have structure available in advance of st…

It's a step forward for sure, but structures change over time to perform their function. The method described here only returns a static structure. Much more research and development is needed to be able to predict the dynamic behavior and interplay with other proteins or RNA.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#112
post #54

So the median accuracy went from ~58% (2018) to 84% (2020) in 2 years? Does 84% == solved? Also, any low hanging frut implications for longevity tech?

100% accuracy is "solved".

Solving the inverse problem would be even more valuable -- given a specific shape (and other biochemical desiderata), what sequence of amino acids would create that protein?

As hard as the protein folding problem is, the inverse problem is harder still. THAT is the one true grail.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#113
post #30

Sometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.

Not to mention the fact that two years ago they took it from 45% to >60%. If they can continue improving, even with an exponential decay in rate of improvement, this is certainly a stunning example of technological disruption.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#114
post #36

This is a big step forward, but the outstanding question as far as to whether or not this is useful for evaluating novel proteins, is going to be how good is the confidence metric at telling the user to trust or not trust the results. You can see from their examples, that AlphaFold is very good but not perfect. I imagine for some proteins it will still give misleading or erroneous results and if you can’t tell when t…

> the outstanding question as far as to whether or not this is useful for evaluating novel proteins That is not an outstanding question. The test on which DeepMind scored high marks is a test of how well the algorithm folds novel proteins -- proteins whose ground-truth structure has not yet been published.

You missed the actual outstanding question in their comment:

> the outstanding question ... is going to be how good is the confidence metric at telling the user to trust or not trust the results.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#115

Title as submitted is hyperbole, please fix?

I agree. "AlphaFold achieves a median score of 87.0 GDT". While this is a major advance, to me 100 GDT would be 'solved', not 87.

By this metric, nothing has been ever solved in natural sciences. So this is not a useful metric.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#117
post #89

Earlier quoted context omitted.

"A few" does appear quite dismissive of the enormous amounts of effort in structural biology so far. There are more than 170,000 structures in the PDB right now. To determine potential targets for drugs we have to understand what the proteins do. Having the structure is not really enough for that, it doesn't tell you the purpose of the protein (though it certainly can give you some hints). In most cases the proteins…

170k is "a few" compared to 180 million (i.e. the size of the PDB as soon as someone runs AlphaFold over everything in the UniProt.) > In most cases the proteins were determined to be interesting by other experiments, and then people decided to try and solve their structure. Yes, that's what we're doing right now , because structure is not a useful predictor, because we don't have structure available in advance of st…

Determining what a protein structure does might be even harder than folding. Right now we can't really do that ab initio, you have determine the activity in the lab and then look at the structure. And that allows you to potentially identify this motif in other proteins.

If someone produces an AI that you give a sequence and it tells you what the protein does exactly, I'd be extremely impressed. I don't see that happening soon.

The specifics matter a lot here. We can often determine rough functions for subdomains by homology alone. But that really doesn't tell you the full story, it only gives you some hints on what that protein actually does.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#118
Two years ago, after DeepMind submitted its first set of predictions to CASP (Critical Assessment of protein Structure Prediction), Mohammed AlQuraishi, an expert in the field, asked, "What just happened?"

https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp...

Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable in experimental measurements), we can confidently answer AlQuraishi's question:

Protein Folding just had its "ImageNet moment."

In hindsight, AlphaFold v1 represented for protein structure prediction in 2018 what AlexNet represented for visual recognition in 2012.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#119
post #105
post #75

Earlier quoted context omitted.

This is about proteins, not DNA.

Proteins which are coded by DNA.

So what? The DNA only codes for the RNA and amino acid sequence. Structure determination is yet another topic. When we determine the protein structure we already know the sequence. Neither DeepMind has to look at the DNA to train their DNN.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#120
post #36

This is a big step forward, but the outstanding question as far as to whether or not this is useful for evaluating novel proteins, is going to be how good is the confidence metric at telling the user to trust or not trust the results. You can see from their examples, that AlphaFold is very good but not perfect. I imagine for some proteins it will still give misleading or erroneous results and if you can’t tell when t…

I was wondering the same thing. But I also wonder if having good guesses makes the x-ray crystallography and other experiments to verify a given protein easier/cheaper/quicker? I don't know enough about the actual techniques to have an informed opinion but I would think it would be helpful.
Post reply on HN