Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

391–400 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#392
post #327

Earlier quoted context omitted.

My understanding is, that it's always 100 new structures, which is a small fraction of the total structures identified in that year. The reason why the top score in one year, can be lower than in the previous year, is that the test (the 100 structures to guess) is always new and different, so it can end up being 'harder' than the year before. Luck will also play a small role. Another explanation for a reduction in th…

Only 100 new structures each test cycle? That seems a very small test set size ... Is it really possible to select 100 new structures which together are likely to represent a meaningful increase in the sample generalization versus the prior years test set ...?

Given that we only know the structure of on the order of 100k proteins, we might only get another 10k new ones per year. I guess.

Using 1% of those (presumably from the more-often-reproduced subset) for this challenge seems reasonable? Note that the structures have to remain secret up until the challenge, and presumably all those teams uncovering the structures don't want to have to wait up to 2 years every time to actually make their results public.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#393
post #7

Earlier quoted context omitted.

> but many were skeptical as to the longevity of the performance DeepMind was able to achieve For a non-biologist, on what is this skepticism based? Just purely based on following ML news it looks like the trend for ML solutions has been that they've overtaken expert-systems once they've gained a solid foodhold in a field. Maybe this is some perception bias. Are there any cases where ML performed decently but then hi…

> Are there any cases where ML performed decently but then hit a ceiling while expert systems kept improving? Yes, this describes entire history of AI including several boom-bust cycles. In particular the 80's come to mind. Yes the practitioners think that there's no technical barriers stopping them from eating the world, but that's exactly what people thought about other so-called revolutionary advances. Although to…

I think the disconnect this time around is in productionization. We're getting breakthroughs in a wide range of problems, and translating those gains in the problem space into 'real' stable, practical solutions we can use in the world is the remaining gap, and often takes years of additional effort. It's still really expensive to launch this stuff, and often requires domain expertise that the ML research team doesn't have.

We're seeing a lot of this pattern: ML Researcher shows up, says 'hey gimme your hardest problem in a nice parseable format' and then knocks a solution out of the park. The ML researcher then goes to the next field of study, leaving (say) the doctors or whatever to try to bridge the gap between the nice competition data and actual medical records. It also turns out that there's a host of closely related but different problems that ALSO need to be solved for the competition problem to really be useful.

I don't think this means that the ML has failed, though; it's probably similar to the situation for accounting software circa 1980: everything was on paper, so using a computerized system was more trouble than it was worth. But today the situation in accounting has completely flipped. Apply N+1 years of consistent effort improving data ecosystems, and the ML might be a lot easier to use on generic real world problems.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#394
post #346

Like this is awesome and a huge advancement but one thing that worries me with an AI solution is that it doesn't really draw us any closer to the why. Why do proteins fold the way they do? We can predict the resulting structure which is extremely significant, we have no clue why. While we get the insight of being able to predict some structures we don't get the insight of why things are happening the way they are. In…

> Why do proteins fold the way they do?

I think the why is pretty clearly understood (https://en.wikipedia.org/wiki/Protein_folding), in the same way that we understand the mechanism behind the three body problem in physics or quantum computing. But that does necessarily imply that there is an efficient way for us to simulate/predict the results of having nature play out those mechanisms.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#395

Earlier quoted context omitted.

Actually, no! Or at least, the budget they used ( In other words, it's less like GPT3 and more like ImageNet.

I don't know - tens of thousands per train is not accessible for most academic institutions when you consider the necessity of ablation studies, experimentation, etc.

For a topic like protein folding, it should be

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#396

Earlier quoted context omitted.

I don't know - tens of thousands per train is not accessible for most academic institutions when you consider the necessity of ablation studies, experimentation, etc.

For a topic like protein folding, it should be

> For a topic like protein folding, it should be

Well, I've worked in some academic deep research labs and they did not have the money to do the experiments they wanted to do.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#397

Earlier quoted context omitted.

Is the cost to train really the relevant metric for developing this? It seems like the salary's involved are probably at least 10x whatever they spent on hardware.

>Is the cost to train really the relevant metric for developing this? Yes, because they must release sufficient information for others to recreate the AI model, according to the rules of entering CASP.

I was replying in the context of the grand parent:

> Did AlphaFold2 also have the biggest budget? :)

And then the parent

> Actually, no! Or at least, the budget they used (I'm not sure how the cost of replicating the model in the future is relevant in this context. We appear to be discussing the cost of developing this model from scratch, such as what it would have taken an alternate team to create and submit this if DeepMind never got involved.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#398

Not knowing a lot about biotechnology, I read the article and it sounds great, but how big is this as a gamechanger? Can someone comment on how big are the implications of this in, let’s say, 5 years from now, on day to day life? Does this mean that biotech is going to explode? Or just that drugs will come to market faster, perhaps cheaper for rare diseases, but from the same industry structure as always?

One young lady I knew worked on neural algos recognition of X-ray images. They always had single digit, bizarre artifacts, where the program can't sometimes recognise the very data it was trained on with most minute differences. Other artifact was that the most "stereotypical cases" were least reliably recognised, and they hot a lot of flak for screwed up live demos, where a radiologist put a very, very obvious tumor…

I find this disingenuous. Yes, its important that the algos can perform well on real world data, but the framing of this post begins with an anecodote about one person who had a bad model, and implicitly extrapolates that these problems are generalized throughout all neural nets.

One could say the same thing about programmers automating a task, or a number of other trivial examples. I would lean towards assuming deep mind has competent model validation teams vs. not, even if data science is hard.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#399
post #325

Earlier quoted context omitted.

> this isn't the typical HN pedantry. Then launches into what can only be recognized as an exercise in pedantry.

"Pedantry" implies that the distinction is not meaningful. This is true if you're only paying attention to how this system can be utilized to answer questions posed to it. This achievement by itself, however, does not do much to push the science of protein folding much further. Those advances will come when people poke, prod, and break the model to develop a unified theory for protein folding.

The "science" of protein folding has a primary goal: to predict the structure of a protein given it's constituent parts.

This is what alphaFold does, and it's been verified to produce results at an apparent accuracy at or above something like X-ray protein crystallography. The advances will come, after these results are validated and accepted by the scientific community as whole, simply when groups start using this technique to immediately access the structure of proteins that in the past would be prohibitively expensive and time consuming or down right impossible to access before, and then use that knowledge to do their work.

You seem to think the first thought a researcher will have after this becomes widely available is, "Oh hey, I can now accurately predict the shape of an arbitrary protein which unlocks untold potential scientific progress on numerous scientific fronts, but the thing I want to spend my time on is trying to replicate the results of the network myself, so I can do it manually thousands of times slower...", which is patently inane.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#400

Earlier quoted context omitted.

In their CASP abstract[1] they mention alternatives to typical co-evolution features which improve performance in shallow MSA depths. [1]: https://predictioncenter.org/casp14/doc/CASP14_Abstracts.pdf...

It doesn't matter so much how they perform the feature extraction, so much as what their inputs to the feature extraction are. This model requires a collection of wild-type proteins in an accurate MSA. Producing an accurate MSA is hard even if you have many homologs. They require protein homologs which means they can "only" do this for wild-type proteins. This work is useless with mutant and synthetic proteins. This…

> Producing an accurate MSA is hard even if you have many homologs.

To assess co-evolutionary couplings the amount of homologs in the MSA is not as important as the number of effective sequences (i.e. sequence depth and diversity) in it.

> They require protein homologs which means they can "only" do this for wild-type proteins.

Even remote homologs work, as shown by the widespread use of HHM-based methods in the prediction pipelines.

> This work is useless with mutant and synthetic proteins.

Unless you generate a flurry of data with them using deep mutational scanning for example. As long as correlated mutations are present in the MSA the technique should work as expected no matter where the protein sequences originated.

Post reply on HN