Live data from Hacker News

AlphaFold: a solution to a 50-year-old grand challenge in biology

deepmind.com

561–570 of 683 posts

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#561
post #471
post #416

Earlier quoted context omitted.

Then we get the really fun question: if the experimentally determined structure is only 90% accurate, can machine learning actually reach 100%? Can you learn exact truth from inexact examples? Which gets into the concept of whether the ML model has actually learned some deeper conceptual ideas than we have, some deeper truth about how this works. If so, can we somehow extract that truth, or is it truly a black box th…

If you have an experimental error that is somewhat normally distributed around the mean, the the AI should, with enough examples, learn what the rules are that are closest to the mean. Because it will minimize the sum of errors. So i do think the results could be more accurate than measurement.

I don’t think we can assume the errors are normally distributed. It’s possible researchers are biased in a particular “direction”, away from 0 on all dimensions of this problem.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#562
post #87

Earlier quoted context omitted.

For-profit corporations that value protein engineering will beat a path to DeepMind's door ASAP, like pharmas. Protein conformation prediction is essential when engineering new small-molecule drug compounds that must 'dock' with the specific proteins that regulate disease. Knowing how to create a protein with the precise shape to become biologically active has soaked up a lot of R&D funding toward pie-in-the-sky tech…

you give pharma too much credit. I had built a previous system to do something similar to this that produced excellent results and tried to give it away for free to Genentech, which ignored me. They said it didn't work for their purchasing department.

I feel that the "produced excellent results" has a lot of unpack there.

It obviously wasn't scoring 90+ in CASP.

Actually, after reading your linked blog post, it's pretty obvious why they weren't exactly chomping on the bit:

"To gain insights into the receptor’s dynamics, Kai performed detailed molecular simulations using hundreds of millions of core hours on Google’s infrastructure, generating hundreds of terabytes of valuable molecular dynamics data."

Hmm, yes, hundreds of millions of hours of cpu time, hundreds of terabytes of data, who says no to that? It doesn't even seem seem to attack generalized protein folding in general. It really seems like the plan was, "let's attack this problem with a Google-sized firehose" rather than created a fundamentally different algorithm that had game-changing results.

Comparing your system to AlphaFold seems like your really bending the truth here.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#563

Earlier quoted context omitted.

I particularly like the rant on pharmaceuticals companies lack of basic research. My impression has been that medical progression have been slow for quite some time, nice to see that there are some truth to that. In the end software and tech companies might just eat up the pharmaceutical industry as well. - It's all just code at some level. The Deepmind team did this with ; "We trained this system on publicly availab…

In a prior(/n) life I worked on Protein folding, and participated in CASP. This was a/the "holy grail" problem of molecular biology, long thought to be an automatic Nobel. It's somewhat unfair to characterise developments prior to this as insignificant. In fact by the time I was working on it, that "automatic Nobel" was no longer assumed, because the field had made quite a bit of progress, in many tiny steps by many…

> How many false start does it take before you do that run with 128TPUs "for a few weeks" that works?

This is a big issue that most people miss. Having easy access to vast computational power makes such a difference for experimentation.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#564
post #520

Earlier quoted context omitted.

> I don’t think we would do ourselves a service by not recognizing that what just happened presents a serious indictment of academic science. Much like other fields, I do begin to question the academic structure to making advances. It appears something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other majo…

> The other major case in point being NLP Speaking of which, Google Translate was published in 2006, but when did the "learning from data" approach became an accepted idea in machine translation? I think the earlier attempts at machine translation were more about trying to codify grammar rules in software, than doing statistical learning from large text corpuses? I remember in 2002, the approach of leaning protein su…

I think Netflix model of simplifying then translating should work better with internet forums and blog posts. Looking at how google works I was hoping someone in google wold adopt it and release a competing product against google translate https://arxiv.org/abs/2005.11197

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#565

Earlier quoted context omitted.

> Are they talking about creating an actual, physical protein in a lab and observing how it folds? Exactly. Researches purify the folded protein and then use methods such as X-ray crystallography, nuclear magnetic resonance, and cryo-electron microscopy to determine its three-dimensional atomic structure.

If even the experimental approach is only 90% accurate, how do they know which 90% is accurate?

I’m not a protein crystallographer, but here’s my generalist take.

We understand the physics of e.g. X-ray diffraction pretty well, so we can fit pretty decent forward models for the x-ray data given a proposed structure. The hardest task here is getting a good enough guess at the structure to optimize the physical model, and it’s my impression that people use an iterative model refinement workflow. At least that’s how it’s done in condensed matter materials.

There are many sources of experimental uncertainty, like the non-ideal nature of the x-ray source and optics, and the fact that the atoms in the protein are not static but have some thermal fluctuations. so at the end of the refinement you still have some uncertainty on your model parameters (the interatomic distances for proteins I guess), but if you are careful you can calibrate these uncertainties pretty well.

This paper looks like a really good detailed discussion: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4080831/

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#566
post #226

Earlier quoted context omitted.

> So it wasn't out of reach for academia, pharmaceuticals, or others with a bit of resources. How much does hiring a deepmind-like team cost though? (massively more than the TPU resources?) Still within reach of pharmaceutical industry I guess, but maybe not so easy for academia.

From what I can gather, Google bought Deepmind for 500 million USD in 2014, they have outstanding debt to its parent company as of 2019 of 1.3 billion USD. And they had income around 100 million in 2019 but it's all against Google, so looks like a 2 billion +/- 0.5 operation so far, and who knows if they pay for compute. Other articles place the runrate at 500 million per year in 2019. Which means 500 million * 6 yea…

Only the energy cost savings google got from Deepmind probably already makes it a very profitable acquisition https://deepmind.com/blog/article/deepmind-ai-reduces-google...

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#567

Earlier quoted context omitted.

> I don’t think we would do ourselves a service by not recognizing that what just happened presents a serious indictment of academic science. Much like other fields, I do begin to question the academic structure to making advances. It appears something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other majo…

Academia keeps employing people who have done well in classes and within fine bounds. Its a careerist track. Industry cares about results, its more meritocratic

> Industry cares about results, its more meritocratic

Industry cares about positive results. If you're not allowed to fail, you will be afraid to explore. That's what Academia is. Then, the industry reaps the fruit of that exploration, which is as it should be.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#568
post #82

Earlier quoted context omitted.

It's an improvement- and a big one- but not a solution to the problem. It mainly shows just how stuck the community had gotten with their techniques and how recently improvements in DNNs and information theory methods can be exploited if you have lots of TPU time.

It’s officially recognized as a solution.

I think that missed the mark, regardless of the rest of the discussion. It's like saying that the winner of the DARPA Grand Challenge for self-driving cars "solved" autonomous driving back in 2010.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#569

Earlier quoted context omitted.

I think the disconnect this time around is in productionization. We're getting breakthroughs in a wide range of problems, and translating those gains in the problem space into 'real' stable, practical solutions we can use in the world is the remaining gap, and often takes years of additional effort. It's still really expensive to launch this stuff, and often requires domain expertise that the ML research team doesn't…

Next time you fly through a busy airport, think about the system which assigns planes to gates in realtime based on a large number of variable factors in order to maximize utilization and minimize waits. This is an expert system design in the 80's and which allowed a huge increase in the number of planes handled per day at the busiest airports. Or when you drive your car, think about the lights-out factory that built…

Well, we can agree that world peace is off the table!

Beyond that, let's notice that expert systems did indeed change how airports and freeways work: They improved the areas where they solved problems. Deployment happened.

What we're seeing now is new classes of previously unsolvable problems falling. Deployment in medicine is known to be particularly hard, but not impossible. My read on the situation is that there have been a number of ML applications in the current round that have been kinda-successful 'in vitro' and failed in deployment. That doesn't mean that all deployments will fail.

Furthermore... Neil Lawrence points out that in most cases we change the world to fit new technologies. For example, mechanized tomato pickers suck, so we develop a more machine-resistant tomato. Cars break easily on dirt roads, so we pave half the planet. ML/AI somehow flips people's expectations of how technology works, and expect the algorithms to adapt to the world. This is almost certainly wrong.

"it's specific problems chosen specifically for their likely ease in solving using current methods. DeepMind didn't decide to take on protein folding at random; they looked around and picked a problem that they thought they could solve."

I'm actually not sure this is at all true. Protein folding is a long-standing grand challenge on which no current methods were working. My guess is that it was initially chosen for potential impact, and chased with more resources after some initial success.

Re: AlphaFold: a solution to a 50-year-old grand challenge in biology

#570

Earlier quoted context omitted.

> I don’t think we would do ourselves a service by not recognizing that what just happened presents a serious indictment of academic science. Much like other fields, I do begin to question the academic structure to making advances. It appears something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other majo…

> something is rotten in the state of academia. Oddly it's academia doing incremental improvements to existing methods but industry making novel leaps and bounds... The other major case in point being NLP You have to realize that corporate research labs had a high level of recognition back in the 20th century. Labs like the Bell Labs, the RCA Laboratories, or the IBM Research, privately-funded, had a reputation that…

> Interestingly, for those labs to exist, being a monopolistic megacorp is a requirement. It appears to me that today's FAANG monopoly allowed the creation of Google Deepmind and OpenAI,

AFAIK OpenAI is still independent, despite its recent closeness with Microsoft. Deepmind existed and was active well before being acquired by Google. All these to examples prove is that today's big, monopolistic corporations tend to acquire research labs, not that they are a requirement for their existence or successful activity.

Post reply on HN