Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

251–260 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#251

Earlier quoted context omitted.

Because only the experts in this field get to tell us, the laymen, what "solving the protein folding problem means", and they defined it not as "perfect" but as "more than good enough to be acceptable as correct result". Which this did. X has actually solved Y. That's not so much "massively cool", that's historical.

I think the “they” you’re referring to is only whatever PR person wrote the headline. Nowhere in the substance of this (PR!) post does it refer to it as anything but a great leap. When an expert in the field outside of deepmind says protein folding has been solved, I’ll believe it.

It does appear other experts in the field are claiming this: https://twitter.com/MoAlQuraishi/status/1333383769861054464

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#252
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?

I think it's that a score of >90 means the result is within the error bars of whatever particular experiment was chosen to be the "reference".

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#253
post #234

Does this obsolete Folding@home?

My question exactly; or Rosetta @ home, or any of the other protein folding "@home"s. I participate in a few, but would gladly donate my compute resources elsewhere if this is no longer necessary.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#254
post #143
post #30

Sometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I don’t think you can say DeepMind could ever be more accurate to the true physical structure since it was built on the same experimental structures that it is being compared to. The limit of accuracy is the experimental data. However, I think we can say that a DeepMind prediction could at least be as good as a new experimental structure.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#255

"AlphaFold achieves a median score of 87.0 GDT". Game changing, and a huge improvement, but not 100% solved. Also this is for static folding. Dynamic folding and interaction is a much harder problem. Those need to be tackled too before I would consider protein folding 'solved'.

It's probably never going to be solved though right. To truly solve protein folding we'd have to have a program that can stimulate a small but still significant system at the QM level; looks like deep learning can get us 60% (conservatively estimating the whole problem domain ) but not all the edge cases, just like it did in other problem domains as well.

Despite this breakthrough by DeepMind, at this point we still do not understand protein folding. That makes it very hard to say precisely which features would be required to do the simulation correctly.

DeepMind/AlphaFold might have something to contribute there too, depending on how interpretable their network model(s?) are.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#256
No, they didn't. They approximated a solution to protein folding.

The two are different concepts -- this isn't the typical HN pedantry.

"Solving" the problem would entail developing an interpretable algorithm for taking a string of amino acids and determining the 3D structure once folded.

Approximating a solution would entail simulating that algorithm, which is what their neural network is doing. It is of course usually accurate, but you would expect this with any suitable universal function approximator.

Props to DeepMind and congrats to CASP but is it not obvious that this is more hype-rhetoric for public consumption?

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#257

Earlier quoted context omitted.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…

This is the cost of training the final architecture with all the refinements enabled by years of research. These years of research involved trying many different architectures, many of which received as much or more compute time than the final system. The price of training the final architecture is meaningless. Researching and training AlphaGo was expensive but it enabled the ideas and development of AlphaZero which…

> The price of training the final architecture is meaningless.

Meaningless in historical terms, but meaningful in future terms. It's meaningless how long the training took because there were countless resources spent to get to that point. It's meaningful in the future, because we know that training times are fairly short, and iteration can be done fairly quickly.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#258

Earlier quoted context omitted.

ML is a super overloaded term. There are definitely cases where machine learned statistical solutions do not perform as well as the systems tuned by the experts, but if you can define the task well and get the data for a deep solution, usually those will overtake.

This. I believe technically just linear regression could be considered "machine learning".

I've seen people at bio conferences actively calling linear regression machine learning.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#259
post #143
post #30

Sometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I'm not a biologist but I'm not sure that follows. It could be that the experimentally-derived structure is 100% accurate to the actual physical structure but getting 90% of your predicted residues to match that is enough to get an accurate prediction of protein behavior and hence "competitive."
Post reply on HN