Assuming optimistic further progress, what are the implications of accurately predicting protein folding? What are we hoping to discover, or succeed in doing?
AlphaFold: a solution to a 50-year-old grand challenge in biology
221–230 of 683 posts
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#222Earlier quoted context omitted.
I expect this to be quickly replicated once published. Training data is public and training compute is not enormous and AlphaFold of 2018 did get replicated.
How do you define enormous? "It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks". Also last time it took about a year for good replications to pop up.
While substantial, the resources here are well within reach of many labs, research institutes, and organizations. For this result this big, I'd guess we'll have 2-6 additional implementations in the next 18 months. The problem has been 'open' for 40+ years, so that's lightening fast!
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#223I hate headlines like “X has solved Y.” How often have we see computer vision and natural language solved at this point, whenever a model does well enough in a benchmark? Their own article doesn’t even have that headline. This is a massively cool thing that’s happened. Why ruin it with a massively hyperbolic headline?
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#224Earlier quoted context omitted.
I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…
This is the cost of training the final architecture with all the refinements enabled by years of research. These years of research involved trying many different architectures, many of which received as much or more compute time than the final system. The price of training the final architecture is meaningless. Researching and training AlphaGo was expensive but it enabled the ideas and development of AlphaZero which…
The research is the giant shoulders you stand on, the compute cost is the price of the tool you need to do the present-day work.
Both are relevant but the shoulder’s of giants are generally more accessible, particularly if we’re talking about published research and not proprietary tech.
A competing team is not starting from the same place the DeepMind team started at 5 or 10 years ago.
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#225Earlier quoted context omitted.
The set of all proteins which can potentially be expressed in an organism is known. Now maybe we also get decent (static) structure information for these. But the interaction of a virus with the host cell is much more complex. There is much more than just an amino acid sequence involved. And these parts are all moving, so a static picture as we now can create faster than before does not contain all the information ne…
Precisely why I referred to it as a different and harder problem
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#226Earlier quoted context omitted.
AlQuraishi described the progress made in CASP13 (2018) as “two CASPs in one”. This one is an even bigger breakthrough.
I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…
How much does hiring a deepmind-like team cost though? (massively more than the TPU resources?)
Still within reach of pharmaceutical industry I guess, but maybe not so easy for academia.
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#227Earlier quoted context omitted.
>The set of all proteins which can potentially be expressed is known. Sure, "known", but it's on the order of 20^10000. It won't fit in the entire visible volume of the universe.
No, the genome of the host is much smaller than the theoretical number of combinations. There are about 20 to 30k different proteins in a human cell (about 20k directly encoded on the DNA).
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#228How will this get into the hands of those who could use it?
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#229Earlier quoted context omitted.
That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?
The threshold for "real" in particle physics is +5 sigma. Which takes a lot of data.
Re: AlphaFold: a solution to a 50-year-old grand challenge in biology
#230Earlier quoted context omitted.
That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.
I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?