Live data from Hacker News

Alphafold

github.com

61–70 of 170 posts

Re: Alphafold

#61
post #38

Honest question: since AlphaFold doesn't really _solve_ the protein folding problem (it's NP-complete after all), but only _approximates_ solutions very well, what are the real impacts of this? Isn't a good approximation of a protein enough to cause unexpected problems? How do we know that an approximate structure will perform the same as the correct solution?

There is a lot of bias in the chat here from a more chemistry and pharma slant. If you ignore this AlphaFold solves in a very meaningful way the problem blocking a lot of science investigation.

For comparative and evolutionary analysis structure is far more conserved than sequence. Especially in things like viruses or anything with a high rate of reproduction like bacteria. Just knowing the general fold or overall structure is enough to do structural alignment and tell if two genes are related on that basis, even if their genomic sequence is completely dissimilar. Large groups of researchers rely on sequence homology built from sequences of known structure.

But AlphaFold works well in new sequence space to far more accuracy than is needed. If we had an AlphaFold prediction for every known sequence suddenly the evolutionary relationships between all genes and even all species would be far clearer. This on its own unlocks a new foundation to reason about function and molecular interaction with a wholistic systems view without gaps in what we can know with some reasonable assurance.

For an analogy think of the difference between having books in different languages describing objects. You know what some of the book in English might say but you dont even know if the book in Spanish is even talking about the same things. AlphaFold is like an AI that transforms all the books into picture books and now we can use image similarity or have one person look at all pictures.

Re: Alphafold

#62
post #23

Earlier quoted context omitted.

> A disturbing thing is that the architecture is much less novel than I originally thought it would be, so this shows perhaps one of the major difficulties was having the resources to try different things on a massive set of multiple alignments. This is something an industrial lab like DeepMind excels at. Whereas universities tend to suck at anything that requires a directed effort of more than a handful of people. Y…

> high heat-to-light ratio Sorry for the ignorance but what does this mean?

It’s trying to say light is more valuable than heat, or some such folksy thing. I cook steak in the dark so I don’t find it to be a very insightful metaphor.

Re: Alphafold

#63
post #23

Earlier quoted context omitted.

> A disturbing thing is that the architecture is much less novel than I originally thought it would be, so this shows perhaps one of the major difficulties was having the resources to try different things on a massive set of multiple alignments. This is something an industrial lab like DeepMind excels at. Whereas universities tend to suck at anything that requires a directed effort of more than a handful of people. Y…

> high heat-to-light ratio Sorry for the ignorance but what does this mean?

Incandescent light bulbs are generally very inefficient in producing light, compared to LED for example. They produce a lot of heat and not much light for which they are made.

So in this context I suppose that gp implies that these threads don't provide much meaningful discussion but rather lots of hand waving.

Re: Alphafold

#64
post #47

Does anyone on HN work in bio or drug discovery? Could you give an overview of how people can leverage this (or how you might?). From reading around about it, it sounds like there's often a need to find a certain type of molecule to activate/inhibit another based on shape and the ability to programmatically solve for this makes the searching way easier. Is this too oversimplified/wrong? How will this be used in pract…

It can be an aid in drug development, and can perhaps assist a bit in tuning small molecule drugs for more stable binding. Though I think the major impacts will be two-fold: (1) The field of structural biology is going to see a change, with much more data available. Some structures of difficult to crystallize proteins will be solved, which may lead to much greater biological understanding. We may enter a time, where…

For those that are unaware, industrial protein design is a multibillion dollar industry. For example, decades ago Genentech and Dow Corning formed a company that developed proteases (proteins that cut other proteins) that worked at much higher temperatures than the ones in nature. This was then sold to P&G and other major laundry companies (laundry detergent contains idle enzymes activated by the heat of the laundry water, and they go clean up. "Protein gets out protein" was the marketing jingle.

That was a few billion dollars right there and almost all the work was done by hand by lab scientists.

Re: Alphafold

#65
post #23

Earlier quoted context omitted.

> A disturbing thing is that the architecture is much less novel than I originally thought it would be, so this shows perhaps one of the major difficulties was having the resources to try different things on a massive set of multiple alignments. This is something an industrial lab like DeepMind excels at. Whereas universities tend to suck at anything that requires a directed effort of more than a handful of people. Y…

> high heat-to-light ratio Sorry for the ignorance but what does this mean?

It's an idiom implying that there's a lot of chatter and bold claims, but very little of it is factual or informative.

Re: Alphafold

#66
> The AlphaFold parameters are made available for non-commercial use only, under the terms of the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. You can find details at: https://creativecommons.org/licenses/by-nc/4.0/legalcode

Does CC BY-NC actually do this? As far as I can tell it only really talks about sharing/reproducing, not using.

Or is the only thing prohibiting other commercial use the words "available for non-commercial use only"?

Re: Alphafold

#67
post #32

Honest question: since AlphaFold doesn't really _solve_ the protein folding problem (it's NP-complete after all), but only _approximates_ solutions very well, what are the real impacts of this? Isn't a good approximation of a protein enough to cause unexpected problems? How do we know that an approximate structure will perform the same as the correct solution?

> (it's NP-complete after all)

Protein folding is a physical/biological phenomenon. AFAIK we don't currently have a proper exact mathematical formulation of the problem that would let one determine its complexity.

You may be referring to this paper [1]. It only claims that one particular optimization problem, believed to give a solution to protein folding problems, is NP-hard. So, even if a suitable exact formulation exists, it is not yet proven that protein folding is hard, although it for sure seems plausible.

By the way, it is perfectly possible today to solve some very large-scale NP-hard problems (think millions of variables and constraints) in reasonable amounts of time (think minutes or hours). Examples are knapsack problems, SAT problems [2], the Traveling Salesman Problem [3] or more generally Mixed Integer Programming [4].

[1] "Complexity of protein folding", 1993, by Aviezri S. Fraenkel

[2] http://www.satcompetition.org

[3] http://www.math.uwaterloo.ca/tsp/

[4] http://plato.asu.edu/bench.html

Re: Alphafold

#68
post #36

Does anyone on HN work in bio or drug discovery? Could you give an overview of how people can leverage this (or how you might?). From reading around about it, it sounds like there's often a need to find a certain type of molecule to activate/inhibit another based on shape and the ability to programmatically solve for this makes the searching way easier. Is this too oversimplified/wrong? How will this be used in pract…

I haven't worked on the drug-side of things, but here my bio perspective: It's kind of out-of-vogue, but consider the "lock and key" model of proteins and small molecules (drugs). For drug design, what you want to do is get a key that fits just one lock (to pull whatever lever) and not others (to avoid side-effects). It's relatively easy to find a molecule that fits a protein, because that protein is what you might s…

First seems reasonable. I've not heard of anything on the later coming even close credibly - though is an obvious holy grail.

Re: Alphafold

#69
I missed an important detail: """an academic team has developed its own protein-prediction tool inspired by AlphaFold 2, which is already gaining popularity with scientists. That system, called RoseTTaFold, performs nearly as well as AlphaFold 2, and is described in a paper in Science paper also published on 15 July"""

One of the things I say about CASP has to be updated. It used to be "2 years after Baker wins CASP, the other advanced teams have duplicated his methods and accuracy, and 4 years after, everything Baker did is now open source and trivially reproducible"

now, it's baker catching up to DeepMind and it took about a year

https://doi.org/10.1126/science.abj8754

Re: Alphafold

#70
post #54

So... is it possible to clone this and turn it into a Folding@Home client? How does it do?

no, it wouldn't make sense to do that. Folding@Home is for ab initio where you don't have any prior info for the structure, this is for homology modelling. F@H probes the dynamics of protein folding, this just makes a static prediction.
Post reply on HN