Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

601–610 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#601

Earlier quoted context omitted.

Well, this relies on the assumption that a God also inherently exists outside of time and space, which is debatable even among religious scholars.

God is simply existence and love. Love exists outside of time and space :)

All the love I've ever seen exists firmly inside of time and space.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#602
post #327

Earlier quoted context omitted.

My understanding is, that it's always 100 new structures, which is a small fraction of the total structures identified in that year. The reason why the top score in one year, can be lower than in the previous year, is that the test (the 100 structures to guess) is always new and different, so it can end up being 'harder' than the year before. Luck will also play a small role. Another explanation for a reduction in th…

Only 100 new structures each test cycle? That seems a very small test set size ... Is it really possible to select 100 new structures which together are likely to represent a meaningful increase in the sample generalization versus the prior years test set ...?

100 structure with 100+ amino acids each, so it's not quite as bad. Part of the folding information is contained within a distance of a few amino acids, while some (the harder part and crux of the problem) is farther away.

But yeah, compared to other fields, the size of training/test sets is sometimes pretty small in ML for life sciences.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#603

Earlier quoted context omitted.

Given that we only know the structure of on the order of 100k proteins, we might only get another 10k new ones per year. I guess. Using 1% of those (presumably from the more-often-reproduced subset) for this challenge seems reasonable? Note that the structures have to remain secret up until the challenge, and presumably all those teams uncovering the structures don't want to have to wait up to 2 years every time to a…

Interesting ... plenty of opportunity then potentially for the 100 samples to have prediction similarity to the set of published discoveries (for expected or unknown reasons)? I suppose it will take a few more years of repetition for the challenge to confirm that the problem has been been solved -- but I wonder if a new version of the contest is going to be needed as well? Maybe the model accuracy is now high enough…

There are different categories of samples, namely FM and TBM targets. FM targets don't have any similarity to known structures. Roughly a quarter were FM targets. I think the more interesting thing to look at is the size of the multiple sequence alignments (MSAs) which is the basis of this and essentially all methods. They seem to do very well with few MSAs, which bodes well for other targets, although there are families of proteins with few MSAs.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#604

Has anyone got any good other references for this? After some of the dodgy experiments related to alpha zero (comparing to purposefully degraded chess systems), I'd love to see some independent analysis.

The article in Science implies that we have independent confirmation of predictions yielding useful results, beyond the challenge itself: > The organizers even worried DeepMind may have been cheating somehow. So Lupas set a special challenge: a membrane protein from a species of archaea, an ancient group of microbes. For 10 years, his research team tried every trick in the book to get an x-ray crystal structure of th…

Thanks, that really is convincing.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#605

What size was the test set? From what I gather, training was done on 170,000 Amino Acids (features) and the resultant protein structure (labels). This is out of 200 million possible proteins. How many examples were in the test set? EDIT: Looks like N=100 for test set: “Entrants get amino acid sequences for about 100 proteins whose structures are not known” https://www.sciencemag.org/news/2020/11/game-has-changed-ai-.…

As usual with ML, I now wonder how “similar” the test set is to the training set, compared to the examples that are neither in the training set, nor the test set: TODO: 200 million - 170,000 training - 100 test ~= 199.8 million proteins

They trained on 170k sequences/ structures/ proteins, each sequence has 10s to 100s or even 1000s amino acids. Structure is much more conserved than sequence. Out of the 100 targets, roughly 1/4th have no similarity to known structures, so there shouldn't be an overlap for those with the training set. They did very well on those targets.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#606
post #553

There's something I don't understand about protein shapes. There are tons of software solutions – on the web and offline – to visualize the shape of proteins from their sequence of amino acids. How do these work then, if we don't know how the atoms might be arranged in space? For example, this[1] is the code for SARS-Cov-2's Spike(S) protein. From what I understand of this page it's pretty short, only ~1,757 proteins…

Al such proteins have been crystallised and we know their shape experimentally.

Thanks for the answer! That explains it. I was looking just at the amino acid sequence and missing a whole lot.

I read the protein folding and X-ray crystallography articles on Wikipedia and they had most of the answers I was looking for. I also saw a request being made by this JavaScript 3D viewer to fetch the PDB (Protein Data Bank) file for the model, which is a text file with tens of thousands of lines describing the coordinates of atoms in space as well as their bonds and other structures. It even has some metadata about the way the data was collected.

For the spike protein linked above: http://files.rcsb.org/view/6X6P.pdb

I find it fascinating that we're even able to scan the 3D structure of molecules with such precision.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#607

Has anyone got any good other references for this? After some of the dodgy experiments related to alpha zero (comparing to purposefully degraded chess systems), I'd love to see some independent analysis.

The "dodgy experiments" (setting the per-turn computation time to a fixed value) in the chess system were only in the pre-print. In the actual publication, they allowed for full time control of the most up-to-date version of stockfish.

I will admit I may not have kept up to date.

Did they also restore the opening and endings, and use the latest Stockfish?

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#608

Earlier quoted context omitted.

"A sufficiently advanced Artificial Intelligence would be indistinguishable from God." (Way Of The Future - AI Church)

That would require the AI to exist outside of time and space.

I think you could actually argue that it does; it just solved a problem in a relatively short amount of time (iirc the folding@home project has been crunching numbers for over a decade and barely got close), and it doesn't occupy 'real' space since it lives on various computers - it could occupy a whole datacenter, or be contained to a single chip, either way it's in a scale that humans themselves can never exist at.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#610
post #570

Earlier quoted context omitted.

> something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other major case in point being NLP You have to realize that corporate research labs had a high level of recognition back in the 20th century. Labs like the Bell Labs, the RCA Laboratories, or the IBM Research, privately-funded, had a reputation that…

> Interestingly, for those labs to exist, being a monopolistic megacorp is a requirement. It appears to me that today's FAANG monopoly allowed the creation of Google Deepmind and OpenAI, AFAIK OpenAI is still independent, despite its recent closeness with Microsoft. Deepmind existed and was active well before being acquired by Google. All these to examples prove is that today's big, monopolistic corporations tend to…

Yeah, but both labs burn hundreds of millions of dollars. Without Google/Microsoft (not to mention Tesla/YC money), they would have died before bringing these kind of results to market.
Post reply on HN