Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

221–230 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#222
post #205

Earlier quoted context omitted.

I expect this to be quickly replicated once published. Training data is public and training compute is not enormous and AlphaFold of 2018 did get replicated.

How do you define enormous? "It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks". Also last time it took about a year for good replications to pop up.

A year is a fast time to replication in many scientific fields.

While substantial, the resources here are well within reach of many labs, research institutes, and organizations. For this result this big, I'd guess we'll have 2-6 additional implementations in the next 18 months. The problem has been 'open' for 40+ years, so that's lightening fast!

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#223

I hate headlines like “X has solved Y.” How often have we see computer vision and natural language solved at this point, whenever a model does well enough in a benchmark? Their own article doesn’t even have that headline. This is a massively cool thing that’s happened. Why ruin it with a massively hyperbolic headline?

The "solved protein folding" part isn't even in the article. It appears to be clickbait editorialization by whoever submitted the link.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#224

Earlier quoted context omitted.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…

This is the cost of training the final architecture with all the refinements enabled by years of research. These years of research involved trying many different architectures, many of which received as much or more compute time than the final system. The price of training the final architecture is meaningless. Researching and training AlphaGo was expensive but it enabled the ideas and development of AlphaZero which…

> The price of training the final architecture is meaningless.

The research is the giant shoulders you stand on, the compute cost is the price of the tool you need to do the present-day work.

Both are relevant but the shoulder’s of giants are generally more accessible, particularly if we’re talking about published research and not proprietary tech.

A competing team is not starting from the same place the DeepMind team started at 5 or 10 years ago.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#225
post #160
post #144

Earlier quoted context omitted.

The set of all proteins which can potentially be expressed in an organism is known. Now maybe we also get decent (static) structure information for these. But the interaction of a virus with the host cell is much more complex. There is much more than just an amino acid sequence involved. And these parts are all moving, so a static picture as we now can create faster than before does not contain all the information ne…

Precisely why I referred to it as a different and harder problem

There are a lot of different harder problems.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#226

Earlier quoted context omitted.

AlQuraishi described the progress made in CASP13 (2018) as “two CASPs in one”. This one is an even bigger breakthrough.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…

> So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources.

How much does hiring a deepmind-like team cost though? (massively more than the TPU resources?)

Still within reach of pharmaceutical industry I guess, but maybe not so easy for academia.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#227
post #208

Earlier quoted context omitted.

>The set of all proteins which can potentially be expressed is known. Sure, "known", but it's on the order of 20^10000. It won't fit in the entire visible volume of the universe.

No, the genome of the host is much smaller than the theoretical number of combinations. There are about 20 to 30k different proteins in a human cell (about 20k directly encoded on the DNA).

If you are designing proteins, you're not limited to those that are already encoded in the host's DNA.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#228

How will this get into the hands of those who could use it?

Realistically speaking, if you are a scientist who could use this and you mailed DeepMind, they will probably run it for free and send you the result. It would be a good PR.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#229
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?

Excellent question. At somepoint, I think the only answer is, "have a bunch of different people run a bunch of experiments on the same protein."

The threshold for "real" in particle physics is +5 sigma. Which takes a lot of data.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#230
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?

And is it even meaningful for DeepMind to score better than experimental results? How are DeepMind’s results scored then?
Post reply on HN