Live data from Hacker News

$4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

arstechnica.com

21–30 of 37 posts

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#21
post #19
post #12

Earlier quoted context omitted.

(Bioinformatician here). Although I think bench work is the most obvious route and the most likely way these problems will be solved, there are in principle some computational ways they could be addressed. If we had good computational models of how perturbations would affect transcription networks, for example, we could predict these "side pathways" that so often occur in humans but not in mice. But you're right, the…

(former bioinformatician here) I agree, the totality of the interactions for a single cell is so many orders of magnitude above what we are capable of currently modeling that I fear these computational approaches are dangerously over-hyped. Having been privy to the state-of-the-art projects in a lab with a ~5,000 node cluster it was still disappointing to see how rough the whole cell modeling approaches were. It's re…

Over-hyped? In my neck of the woods these approaches are treated with extreme skepticism for just the reasons you mention.

For instance: De novo protein folding is not a solved problem, so how can a simulation predict dynamics for a protein whose conformation isn't even known?

I'm sure in the year 2150 when my grandchildren go to the doctor to be scanned by the tricorder, the results will go to the full-cell (and full-body) simulator...but for now, I think bioinformatics is better served by sticking to higher levels of abstraction like transcript and protein counts (for disease modeling purposes).

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#22
post #9

While it's nice from a technical perspective, this is unlikely to lead to a cancer cure. Having worked in cancer drug development, I can tell you, there is no shortage of cancer targets. Researchers have a list of targets they want to hit, and chemists are pretty darn good at designing small molecule compounds to hit them. The problem in cancer is not that we don't understand individual proteins or the way that drugs…

I agree mostly with your post but there are computational approaches being developed that may help improve our understanding of these cancer networks. They will in no way eliminates or reduce the need for wet lab biology but hopefully it will couple with improvements in high throughput experimental technology to help us design and make sense of experiments targeted at understanding the whole phenomena.

I am well aware of these approaches, having worked on some of them myself in prior work I've done. You always run into limitations of what was known about the biology and how various things interact. The network models will probably work some day, but my point above was that this will only happen after a lot of hard wet bench work happens. There's just too much we don't know right now to use these techniques to develop deep understanding. Sometimes, when we were lucky, they would support an existing hypothesis. But that was only for activity in a single cancer cell line in a specific experiment, not a whole organism which is what matters for a drug.

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#23
post #7

And yet, no results from their massive computation. Schrodinger is well known for being a company of liars and frauds, and unfriendly to open source and other ideals of our community.

To be fair, Gaussian is orders of magnitude worse. See " rel="nofollow">http://www.bannedbygaussian.org/>.

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#25
post #9

While it's nice from a technical perspective, this is unlikely to lead to a cancer cure. Having worked in cancer drug development, I can tell you, there is no shortage of cancer targets. Researchers have a list of targets they want to hit, and chemists are pretty darn good at designing small molecule compounds to hit them. The problem in cancer is not that we don't understand individual proteins or the way that drugs…

(Naive layperson here). Could the manual microsocpes-and-pipets work being done by lab biologists be mechanized, so that you're generating drug candidates in software, testing them in living cells, and using automatically-gathered observations to generate new candidates?

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#26
Wow. This press release even got covered in the NYT!

http://bits.blogs.nytimes.com/2012/04/19/supercomputing-rent...

The actual feat accomplished here would not surprise anyone involved in supercomputing, and companies have done very similar things with non-scientific tasks, such as this article from 2008:

http://open.blogs.nytimes.com/2008/05/21/the-new-york-times-...

describing how the New York Times itself did a similar computation.

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#27
post #9

While it's nice from a technical perspective, this is unlikely to lead to a cancer cure. Having worked in cancer drug development, I can tell you, there is no shortage of cancer targets. Researchers have a list of targets they want to hit, and chemists are pretty darn good at designing small molecule compounds to hit them. The problem in cancer is not that we don't understand individual proteins or the way that drugs…

(Naive layperson here). Could the manual microsocpes-and-pipets work being done by lab biologists be mechanized, so that you're generating drug candidates in software, testing them in living cells, and using automatically-gathered observations to generate new candidates?

It is already done to an extent (high-throuhput machines), however, there's still a lot of old timers in biology who spent too long pipetting and haven't invested yet.

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#28
post #8

Two things struck me about that article, the pharmaceutical guy who felt like with enough cores they could find the cure to cancer, and Cycle's challenge of moving past 50,000 cores. What struck me is that I wonder why Google (or Amazon) hasn't put out the cure for cancer. The actual extent of Google's infrastructure is classified, but using open sources its clear that putting together even half a million 'cores' is…

Google does invest in life sciences startups: http://www.googleventures.com/portfolio#life-sciences

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#29
post #28
post #8

Two things struck me about that article, the pharmaceutical guy who felt like with enough cores they could find the cure to cancer, and Cycle's challenge of moving past 50,000 cores. What struck me is that I wonder why Google (or Amazon) hasn't put out the cure for cancer. The actual extent of Google's infrastructure is classified, but using open sources its clear that putting together even half a million 'cores' is…

Google does invest in life sciences startups: http://www.googleventures.com/portfolio#life-sciences

Perhaps I should be a bit more clear and less snarky. I think Cycle has had a great press release, it has advertised their product well, and like any good press release it doesn't so much read like an advertisement for a particular company as it does as real news. This defines great execution of the PR technique known as 'article placement.'

One of the qualities of making it a great soundbite is that Schrodinger, the company that it nominally the topic of the story, goes on about how their 1,500 core cluster can't give them the resolution they need but a 50,000 core cluster makes everything clear. Understand that a 'westmere' processor is potentially 12 cores (if you use Intel's defintion of threads and I'm sure they do), and the typical motherboard is 2 CPUs so that is 24 'cores' per machine [1]. A 1,500 core cluster is 62 machines, that is a couple of cabinets worth if you're using Supermicro boxes, less than a cabinet if your using OpenCompute type cabinets [2]. And at maybe $3K each that is an investment of $186K, maybe $250K if you included switches. Which is about 1/3 what it would have cost a pharmaceutical company to buy a VAX minicomputer back in the day.

My point is that if you're in a multi-billion dollar market place, you can afford to spend more on your hardware. And even at approximately $5K/hr a 50,000 node cluster is only 2100 Open Compute servers in 24 of their 'triplet' cabinets.

That is about 1MW critical kW of compute power. (1MW being the power commitment you would have to buy from a colocation center to power it) and even that is a fairly small foot print at Amazon, Google, Facebook, or even Apple with its new $1B data center [3].

So when I read the story, I was left thinking "Gee, if using a cluster that was 33x bigger got them such great results, why not use one that 3000x bigger? Wouldn't that just answer the question?" And of course I took a moment to analyze that thought and asked the question every critical reader has to ask which is, "What is this article trying to say anyway? And do I believe it?" And that was when it becomes obvious what the article is saying is that Cycle, the company that makes a living creating virtual super computers for embarrassingly parallel problems out of EC-2 instances has reached the point where they can get 50K cores running the same problem." Which is great and all but like a long story that is just a setup for bad pun, it leaves me feeling jaded, and hence my snarky remark that if all it takes is more cores, Google should stop trying to be a great advertising company and switch to being a pharmaceutical company. Which, when you say that out loud you realize it couldn't possibly be that easy and yes, its a snarky way of expressing irritation that I was lead to believe there was something newsworthy here when there wasn't.

[1] http://opencompute.org/projects/intel-motherboard/

[2] http://opencompute.org/projects/triplet-racks/

[3] http://gigaom.com/apple/apples-new-north-carolina-data-cente...

Re: $4,829-per-hour supercomputer built on Amazon cloud to fuel cancer research

#30

Earlier quoted context omitted.

(Naive layperson here). Could the manual microsocpes-and-pipets work being done by lab biologists be mechanized, so that you're generating drug candidates in software, testing them in living cells, and using automatically-gathered observations to generate new candidates?

It is already done to an extent (high-throuhput machines), however, there's still a lot of old timers in biology who spent too long pipetting and haven't invested yet.

It would be good even for not-old-timers. I know a couple of bioinformatics people and they're definitely experts on one topic: RSI.

Unfortunately not everything is done on a massive scale so not everything is automated.

Post reply on HN