Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

411–420 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#411

Earlier quoted context omitted.

This is word salad.

Rude. I would appreciate substantive criticism, especially when I'm linking papers in Nature starting to do exactly what I'm talking about.

I cannot give constructive feedback to something which is incomprehensible.

"the genome is a sequence that can be language modeled/auto-regressed for depth of understanding by the network"

The genome is not a sequence so much as a discrete set of genes which are themselves sequences which specify construction plans for proteins. That distinction is important.

Language modeling in the context of machine learning typically means NLP methods. Genetics is nothing like natural language.

Auto-regression is using (typically time series) information to predict the next codon. This makes very little sense in the context of genetics since, again, the genetic code is not an information carrying medium in the same sense as human language. Being able to predict the next codon tells you zilch in terms of useable information.

"Depth of understanding by the network" ... what does that even mean???

The above sentence is a bunch of popular technical jargon from an unrelated field thrown together in a nonsensical way. AKA word salad.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#412
Fascinating! AlphaFold (and other competitors) seem to use MSA (Multiple Sequence Aligment) and this (brilliant) idea of co-evolving residues to build an initial graph of sections of protein chain that are likely proximal. This seems like a useful trick for predicting existing biological structures (i.e. ones that evolved) from genomic data. I wonder (as very much a non-biologist), do MSA-based approaches also help understand "first-principles" folding physics any better? and to what degree? If I write a random genetic sequence (think drug discovery) that has many aligned sequences, without the strong assumption of co-evolution at my disposal, there does not seem any good reason for the aligned sequences to also be proximal. Please pardon my admittedly deep knowledge gaps.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#413
What happens when AI is better at everything measurable than humans?

Better at conversation. Better at making people laugh, and generate attraction or other emotions, better at motivating them, and organizing movements, etc.

Clearly we are not ready for such an efficient system... it would be a big disruption to all human organizations and relations. It would start with Twitter botnets and directing sentiment.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#414

Earlier quoted context omitted.

> Producing an accurate MSA is hard even if you have many homologs. To assess co-evolutionary couplings the amount of homologs in the MSA is not as important as the number of effective sequences (i.e. sequence depth and diversity) in it. > They require protein homologs which means they can "only" do this for wild-type proteins. Even remote homologs work, as shown by the widespread use of HHM-based methods in the pred…

I'm honestly not familiar with "deep mutational scanning." Can you share a link? I'm first author on papers related to the structural biology of coevolution and I competed in CASP about a decade ago, but I haven't kept up much since then.

Sure! Here's a paper about the method: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4410700/

And another one about its application in structure prediction: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7295002/

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#415
post #89

Earlier quoted context omitted.

170k is "a few" compared to 180 million (i.e. the size of the PDB as soon as someone runs AlphaFold over everything in the UniProt.) > In most cases the proteins were determined to be interesting by other experiments, and then people decided to try and solve their structure. Yes, that's what we're doing right now , because structure is not a useful predictor, because we don't have structure available in advance of st…

Determining what a protein structure does might be even harder than folding. Right now we can't really do that ab initio, you have determine the activity in the lab and then look at the structure. And that allows you to potentially identify this motif in other proteins. If someone produces an AI that you give a sequence and it tells you what the protein does exactly, I'd be extremely impressed. I don't see that happe…

Five years ago, I would have said the following:

"If someone produces an AI that you give a sequence and it tells you the protein conformation, I'd be extremely impressed".

Sure there are many more things to solve in this space; but that doesn't take away that this is an impressive achievement and does unlock quite a few things (including making more tractable the problem you just brought up). I'm excited to see what DeepMind works on now and what the new state of the world will be just five years from now.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#416
post #143
post #30

Sometimes announcements like this are a bit over-the-top. But what really, to me, cements the 'big-deal' of this is the "Median Free-Modelling Accuracy" graph half way down the page. Scores of 30-45 for 15 years. Now scores of 87-92. This isn't a minor improvement, it's a leap forward.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

Then we get the really fun question: if the experimentally determined structure is only 90% accurate, can machine learning actually reach 100%? Can you learn exact truth from inexact examples?

Which gets into the concept of whether the ML model has actually learned some deeper conceptual ideas than we have, some deeper truth about how this works. If so, can we somehow extract that truth, or is it truly a black box that does the thing we want?

I'm reminded of a sci-fi book I read long ago in which humans are discussing the fact that the science they are utilizing is beyond the scope of a human mind to comprehend- only the AIs can intuitively deal with 12-dimensional manifolds (or something to that extent). Maybe we've reached the doorstep of that future.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#417
post #143

Earlier quoted context omitted.

That is an impressive improvement, but I think you've missed the most important point: >a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure.

I have a related question about this. If experimental methods produce results around a score of 90, what is the baseline we are comparing the DeepMind results against? If the experimental error is equal to the observed DeepMind error, how can we say which one is actually more erroneous?

The "experiments" here use X-Ray Crystallography. Like most methods of measuring anything, we have a pretty good idea of its accuracy under various conditions.

Think of it like satellite imagery of a tree: A score of zero would be a single green-ish pixel, while a score of 100 would show each leaf within the range it naturally moves in due to wind etc. (proteins tend to wiggle quite a bit under natural conditions, as well)

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#418
post #343
post #341

Earlier quoted context omitted.

That's a good point; this system certainly didn't come from nowhere! The protein datasets they used also mostly came out of various NIH-funded projects. What I meant to focus on was that I think DeepMind has less of a pure money/scale advantage in this area than in some others. In something like Go or Atari game-playing, there are many academic groups researching similar things, but their resources are laughably smal…

Personally I think a major part of the secret sauce is Google's internal compute infrastructure. When I was an academic, 50% of my time went to building infra to do my science. At Google, petabytes of storage, millions of cores, algorithms, and brains were all easily tappable within a common software repo and cluster infrastructure. That immediately translates to higher scientific productivity.

You hit the nail on the head here.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#419

Earlier quoted context omitted.

Rude. I would appreciate substantive criticism, especially when I'm linking papers in Nature starting to do exactly what I'm talking about.

I cannot give constructive feedback to something which is incomprehensible. "the genome is a sequence that can be language modeled/auto-regressed for depth of understanding by the network" The genome is not a sequence so much as a discrete set of genes which are themselves sequences which specify construction plans for proteins. That distinction is important. Language modeling in the context of machine learning typic…

> The genome is not a sequence so much as a discrete set of genes which are themselves sequences which specify construction plans for proteins. That distinction is important.

aka a sequence. "a book is not a sequence so much as a discrete set of chapters which are themselves sequences of paragraphs which are themselves sequences of sentences" -> still a sequence

these techniques are already being used, such as in the paper I just linked.

> Being able to predict the next codon tells you zilch in terms of useable information.

You have absolutely no way of knowing that apriori. And autogressive tasks can be more sophisticated than just next codon.

> bunch of popular technical jargon from an unrelated field thrown together in a nonsensical way

Okay, feel free to think that.

There's always this assumption of it "will never work on my field." I've done work on NLP and on proteins and read others' work on genetics. I think you will end up being surprised, although it might take a few years.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#420
post #345
post #336

Earlier quoted context omitted.

"So DeepMind is to the point where it's a question of whether their generated model or the experimentally determined structure is closest to the actual physical structure." While this is an accomplishment, nobody is going to be confusing these models for structures produced experimentally. The CASP metric is for backbone atoms. To have a useful model of protein structure, you really need to have the positions of the…

So it's a really good start, but nobody is going to be throwing these structures into molecular docking simulations for drug discovery or etc just yet. But hopefully those details can be worked out soon enough.

Yeah, there's a huge difference between a 1Å all-atom RMSD structure, and a 1Å backbone RMSD structure. The non-backbone atoms in a protein make up most of the mass and volume. When structural biologists talk about RMSD, this is what they mean.
Post reply on HN