Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

301–310 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#301

Been out of the field for a while, could someone currently in it qualify these results? Hyperbolic title notwithstanding, they approach 90% median free modeling accuracy. The "other 90%" still remains to be solved...

The method relies on multiple-sequence-alignment (MSA) of homologous proteins. This cannot fold arbitrary proteins, only biologically relevant ones that have high quality MSAs available. It's also worth pointing out that the gold-standard for validating MSAs relies on PDBs of folded proteins. This is exciting work that will assist NMR and XRay crystallographers, but it's not a panacea of protein folding. https://gith…

In their CASP abstract[1] they mention alternatives to typical co-evolution features which improve performance in shallow MSA depths.

[1]: https://predictioncenter.org/casp14/doc/CASP14_Abstracts.pdf...

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#302

Just to add to this whole "It's not solved! Yes it is!" discussion. Note that >According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods. So if we go by >= 90 as solved: >In the results from the 14th CASP assessment, released today, our latest AlphaFold system achieves a median score of 92.4 GDT overall across all targets. they so…

[deleted]

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#303

Earlier quoted context omitted.

I mean, credit where credit is due. Google employs some of the greatest names in artificial intelligence and the DeepMind team had a huge chunk of them working on this problem. While the resources may have been available, I don’t think any other single institution had the level of brain power.

It also makes one reconsider the notion that monopolies are entirely bad. This essentially appears to be a vanity project for Google. Though of course they'll benefit from it in many ways, but it's not like they're doing this as the core product of their service. It's a pretty awesome achievement.

It's kind of like a modern day Bell Labs where they have so much excess profit from adtech that they can fund lots of "basic research" or the computer science equivalent of that.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#304
post #139

Title as submitted is hyperbole, please fix?

I agree. If a newspaper published a headline "Dr. Whatever cured cancer (... in some of her patients)" we would find it misleading.

If there was a headline, "Company X with Product Y cured cancer" and it turns out that product Y actually only cured 90% of cancer, I'm pretty sure most people would be happy the headline.

Oh, and to be a true parallel example, in this case the remaining 10% of cancers might not even be cancers, as experimental accuracy of protein structures is only ~90% accurate, the model could very well be more accurate than our current ability to experimentally detect protein structure.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#305
post #30

Sometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.

Why is the graph not monotonically increasing? Does the complexity of the problem to be solved increase each time? If so, does that make the relative improvement from the previous result even more impressive?

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#306

Just to add to this whole "It's not solved! Yes it is!" discussion. Note that >According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods. So if we go by >= 90 as solved: >In the results from the 14th CASP assessment, released today, our latest AlphaFold system achieves a median score of 92.4 GDT overall across all targets. they so…

What do you mean you're not sure how useful ">= 90" is as a criteria?

You literally said why it is useful in your comment:

> 90 GDT is informally considered to be competitive with results obtained from experimental methods.

It's informal because we don't have a true "gold-standard" for determining a protein's folded structure – the best we have is experimental methods of trying to determine the structure which still have a great deal of error (compared to other things we can measure).

So all we can do is say "the GDT between two experimental measurements (of the same protein) is often around 90, so if we get there with predictive models that's pretty much just as good".

As soon as we have better experimental methods for determining protein tertiary structure, you can be sure we will require predictive models to deliver better results too. Until then, the point is that the delta between any two experimental determinations of folded structure is approximately the same as the delta between an experimental determination and an AlphaFold guess. So the AlphaFold guess may as well be an experimental measurement. Except the AlphaFold guess happens fairly trivially (once you give it the DNA sequence[1]), where as the experimental method is involved and expensive.

[1] Or the primary structure, I'm unsure what inputs are given to AlphaFold.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#307
We indeed stand on the shoulders of a small number of giants! I'm infinitely thankful for the work DeepMind is doing. Lets maybe celebrate this accomplishment for one day and start being worried about big tech again tomorrow. Many of the comments here usually suggest that we should live in worries and fear but to my knowledge there is not too much historical evidence for these kind of companies turning evil.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#308

Earlier quoted context omitted.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…

This is the cost of training the final architecture with all the refinements enabled by years of research. These years of research involved trying many different architectures, many of which received as much or more compute time than the final system. The price of training the final architecture is meaningless. Researching and training AlphaGo was expensive but it enabled the ideas and development of AlphaZero which…

Even if you try to account for the overall R&D cost, DeepMind isn't that large an organization by the standards of biomedical research. It's very big and well funded for a computer science research organization, yes, and most CS departments can't match its resources. But the NIH budget is $40 billion, and private pharmaceutical companies do another $80 billion in annual R&D. It's interesting that this kind of breakthrough didn't come from those sectors.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#309
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I don’t think you can say DeepMind could ever be more accurate to the true physical structure since it was built on the same experimental structures that it is being compared to. The limit of accuracy is the experimental data. However, I think we can say that a DeepMind prediction could at least be as good as a new experimental structure.

This seems like an obvious assumption to make, but it isnt always true. It is easier to see why if you are measuring a single value multiple times in order to get a more accurate estimate of the true value. In that case your "model" is simply the mean of all measurements made and can exceed the accuracy of a single measurement.

In this case, the model is predicting values of multiple structures, but patterns could still theoretically be found which allow for predictions beyond the accuracy of a single measurement.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#310

Earlier quoted context omitted.

The distinction you're making between "solved" and "closely approximated" makes logical sense to me. However, if I'm interpreting the AlphaFold results correctly, this distinction isn't practically significant, right? If you can approximate an algorithm with error that is "below the threshold that is considered acceptable in experimental measurements" (to quote another HN comment), then you have something as good as…

It might be the case that the relevant, practical threshold now tightens. For example, perhaps it is easier to experimentally verify a protein shape predicted by an algorithm than it is to experimentally determine the protein shape?

exactly. Even an incomplete map with somewhat limited resolution makes navigation a hell of a lot easier than flying blind. This effectively is a data reduction solution-- if you have a fuzzy shape of the thing you are trying to model, and you learn the mechanics better with each thing you model, your ability to quickly and accurately reach a goal improves
Post reply on HN