Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

171–180 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#171
post #118

Two years ago, after DeepMind submitted its first set of predictions to CASP (Critical Assessment of protein Structure Prediction), Mohammed AlQuraishi, an expert in the field, asked, "What just happened?" https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp... Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable i…

AlQuraishi described the progress made in CASP13 (2018) as “two CASPs in one”. This one is an even bigger breakthrough.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#172

Not knowing a lot about biotechnology, I read the article and it sounds great, but how big is this as a gamechanger? Can someone comment on how big are the implications of this in, let’s say, 5 years from now, on day to day life? Does this mean that biotech is going to explode? Or just that drugs will come to market faster, perhaps cheaper for rare diseases, but from the same industry structure as always?

Getting from DNA structure from tissue samples is relatively straight forward. DNA -> RNA -> unfolded protein is basically one-to-one mapping in most cases. How protein functions depends on how it folds into itself. Once you solve protein folding, you can take DNA sample and see the structure of the molecule without working in lab using crystallography techniques.

Solving protein folding is huge, Nobel in chemistry scale achievement. It would be massive leap for biochemistry.

It seems that Deep Mind solved competition benchmark and made huge leap, but it's just partial solution that works on limited set.

After you have solved protein folding, there is still problem of solving chemical interactions between molecules accurately. Quantum chemistry is extremely compute intensive.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#173
post #118

Two years ago, after DeepMind submitted its first set of predictions to CASP (Critical Assessment of protein Structure Prediction), Mohammed AlQuraishi, an expert in the field, asked, "What just happened?" https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp... Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable i…

How far does the similarity extend? Specifically, the big question for me is whether AlphaFold will be freely available like ImageNet, or proprietary.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#174
post #155

I am puzzled me about “AI-knowledge”. Have we really learnt anything? Is distilling the knowledge from AlphaFold just as a hard problem as solving protein folding?

If you forgot how to do long division, but still had a calculator, wouldn't the calculator still be useful?

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#175
post #115

Earlier quoted context omitted.

By this metric, nothing has been ever solved in natural sciences. So this is not a useful metric.

Has it not? Neuton's laws of motion and Ohm's law are pretty om point

If you can explain how gravity works in a quantum level you'd deserve a Nobel. It's not 100% solved, Newton's Laws of Motion are a model, not a solution. Just like the vast majority of science.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#176

I hate headlines like “X has solved Y.” How often have we see computer vision and natural language solved at this point, whenever a model does well enough in a benchmark? Their own article doesn’t even have that headline. This is a massively cool thing that’s happened. Why ruin it with a massively hyperbolic headline?

Because only the experts in this field get to tell us, the laymen, what "solving the protein folding problem means", and they defined it not as "perfect" but as "more than good enough to be acceptable as correct result". Which this did.

X has actually solved Y. That's not so much "massively cool", that's historical.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#177
post #37

Earlier quoted context omitted.

We can already determine how a few proteins (170k — which sounds like a lot, but which is only 0.09% of all currently-catalogued protein sequences) fold by experimental work. What an accurate model of protein folding allows us to do, is to take our big database of DNA, predict protein foldings for all of it, and then stand up a search index for this database, keying each amino-acid "row" by the "words" of its predict…

Considering that this system "uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks" to determine a single protein structure, making predictions for all proteins encoded in a human genome seems impractical at this stage. With luck, this advance will help lead to discovery and definition of new folding rules and optimizations that will make protein folding predictions for the wh…

Still much faster than synthesizing the protein and then doing NMR or cristallography to solve the structure puzzle what easily takes half a year or more (and very expensive equipment).

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#178
post #87

Earlier quoted context omitted.

For-profit corporations that value protein engineering will beat a path to DeepMind's door ASAP, like pharmas. Protein conformation prediction is essential when engineering new small-molecule drug compounds that must 'dock' with the specific proteins that regulate disease. Knowing how to create a protein with the precise shape to become biologically active has soaked up a lot of R&D funding toward pie-in-the-sky tech…

you give pharma too much credit. I had built a previous system to do something similar to this that produced excellent results and tried to give it away for free to Genentech, which ignored me. They said it didn't work for their purchasing department.

I don't believe you, but I look forward to you showing proof of this with some links (and if you tried giving it for free, I assume you just open sourced the whole deal, so I look forward to a repo link or the like).

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#179
post #118

Two years ago, after DeepMind submitted its first set of predictions to CASP (Critical Assessment of protein Structure Prediction), Mohammed AlQuraishi, an expert in the field, asked, "What just happened?" https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp... Now that the problem of static protein structure prediction has been solved (prediction errors are below the threshold that is considered acceptable i…

AlQuraishi described the progress made in CASP13 (2018) as “two CASPs in one”. This one is an even bigger breakthrough.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that.

In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level.

The Deepmind team did this with ;

"We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 128 TPUv3 cores (roughly equivalent to ~100-200 GPUs) run over a few weeks, which is a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today."

So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#180

"AlphaFold achieves a median score of 87.0 GDT". Game changing, and a huge improvement, but not 100% solved. Also this is for static folding. Dynamic folding and interaction is a much harder problem. Those need to be tackled too before I would consider protein folding 'solved'.

It's probably never going to be solved though right. To truly solve protein folding we'd have to have a program that can stimulate a small but still significant system at the QM level; looks like deep learning can get us 60% (conservatively estimating the whole problem domain ) but not all the edge cases, just like it did in other problem domains as well.

It remains unclear whether QM is required to fold proteins accurately. So far classical methods have shown they require far less computer power to get far closer to the right structure.
Post reply on HN