Live data from Hacker News

A farewell to bioinformatics (2012)

madhadron.com

161–170 of 179 posts

Re: A farewell to bioinformatics (2012)

#161

I have some experience working at a genomics research company and I'll broadly +1 Fred's experience about the industry, although in less negative terms. I got out before I got jaded, so my perspective is a bit more "oh, that's a shame" than his. I really like genetics, bioinformatics, hardware, deep-science, and all that but the timing and fit wasn't right. The tools are written by (in my experience) very smart bioin…

Hrrm, I was in a genetics research lab myself and got annoyed at the inefficiencies myself. In particular, I got frustrating to write & use in-house scripts to run pipelines for compute clusters and then not know what the state of the execution is, where the files are, etc. It's sort of a meta-problem, but I decided to do a startup based on writing good software w/ a good UI to make the problem better (problem = running, monitoring, managing pipelines on clusters that have job schedulers like Grid Engine): http://www.palmyrasoftware.com/workflowcommander/

Maybe it's still in the early going, but I do see how it's going to be real difficult making a living doing this. OTOH, companies like CLC Bio seem like they're doing well for themselves...

Re: A farewell to bioinformatics (2012)

#162
I spent a year in a bioinformatics PhD program and got the feeling I was studying to be science's version of the business analyst. Not knowing enough about the biology or computation, but expected to speak the language of both. And what would my research consist of in such an applied science? Luckily I had another opportunity and became a software developer (which I'm happy with). The worst thing about the experience was listening to so many research presentations where I could tell the presenter didn't understand the science and could barely explain it.

Re: A farewell to bioinformatics (2012)

#163

I have some experience working at a genomics research company and I'll broadly +1 Fred's experience about the industry, although in less negative terms. I got out before I got jaded, so my perspective is a bit more "oh, that's a shame" than his. I really like genetics, bioinformatics, hardware, deep-science, and all that but the timing and fit wasn't right. The tools are written by (in my experience) very smart bioin…

Why wouldn't anyone buy your product? If it is easy to use, and SPEEDS UP RESEARCH TIME, your researcher/PI who is spending thousands on computing clusters will buy your software for their graduate students. Hell, my PI keeps asking me if I need a faster computer so I can run Matlab better/quicker. Really, if I had a software that helped me perform research faster/better/quicker and compare my results to ground truth or gold-standards, that is a much more useful tool than a bunch of hardware for my research. You push out papers fast.

So I disagree with you on your very last sentence (agree with the rest)

Re: A farewell to bioinformatics (2012)

#164
I agree that a lot of effort that is put into bioinformatics is wasted. But it's silly to say that bioinformatics hasn't contributed much to science, and naive to think that dysfunctional software development is less widespread outside of bioinformatics.

Re: A farewell to bioinformatics (2012)

#165
post #131
post #100

Earlier quoted context omitted.

Research is highly competitive business mixed with industry involvement (or government involvement). You have to publish and fast. You have to develop your discoveries into something that can be monetized. You have to collaborate with industry to get funded. You have to cut costs to keep doing what you want to do. And so on. The idea of freedom in (fundamental) research seems long dead. How I long for the freedom in…

> The idea of freedom in (fundamental) research seems long dead Is it so in the US? Or where? Here in Russia it is far from true, at least in the top institutes. As long as you produce publishable results, you may do virtually whatever you want, and nowadays pretty much anything is publishable. And this way you get funding, too, because the funding agency doesn't seem to want you to solve some particular problem, it…

Here in the Netherlands it is. I assumed it to be the same in the Western world, but those kind of generalizations often turn around to bite me in the ass. We, researcher in the Netherlands, have to produce as funding depends on it. Furthermore, as the government funds less and less, we have to get more funding from industry. And finally we have to try to market our research more. This all means that we can not afford to just do whatever we think is best for the only purpose of extending our knowledge. We have to think about our career and the sustainability of our research (strand) in the long run.

That does not mean that we're just lapdogs for industry or Mammon, but it does mean that we're selective in what we do and how we do it.

Re: A farewell to bioinformatics (2012)

#166

Earlier quoted context omitted.

On your first point, my post detailed that 23andMe confirmed it was a GATK bug that introduced the bogus variants and the bug was fixed in the next minor release of the software. There are comments on the post from members of 23andMe and the GATK team that go into more details as well. On your second point. 23andMe had every incentive to pay attention to their output, but it is fair to say it's their responsibility f…

> So what I actually argue in the post (and should have stated more clearly in my summary here) was that GATK is incentivised, as an academic research tool, to quickly advance their set of features with the cost of bugs being introduced (and hopefully squashed) along the way. Sure, I agree with that. And I would agree if you would say "Using bleeding-edge nightly builds of %s for production-level clinical work is a b…

They key is 23andMe was not using bleeding-edge nightly builds but official "upgrade-recommended" releases.

GATK currently has no concept of a "stable" branch of their repo (Appistry is going to provide quarterly releases in the future, which is great).

The flag I am raising is that a "stable" release is needed before it get's integrated into a clinical pipeline. Because the Broad's reputation is so high, it is important to raise this flag as otherwise researchers and even clinical bioinformaticians assume choosing the latest release of GATK for their black-box variant caller is as safe as an IT manager choosing IBM.

Re: A farewell to bioinformatics (2012)

#167
post #163

I have some experience working at a genomics research company and I'll broadly +1 Fred's experience about the industry, although in less negative terms. I got out before I got jaded, so my perspective is a bit more "oh, that's a shame" than his. I really like genetics, bioinformatics, hardware, deep-science, and all that but the timing and fit wasn't right. The tools are written by (in my experience) very smart bioin…

Why wouldn't anyone buy your product? If it is easy to use, and SPEEDS UP RESEARCH TIME, your researcher/PI who is spending thousands on computing clusters will buy your software for their graduate students. Hell, my PI keeps asking me if I need a faster computer so I can run Matlab better/quicker. Really, if I had a software that helped me perform research faster/better/quicker and compare my results to ground truth…

Ahh the efficiency argument.

The trick is, academics often have excess manpower capacity in the form of grad students and post-docs. Even though personell is usually one of the highest expenses on any given grant, they often don't look at ways to improve the efficiency of their research man-hours.

That's not a blank rule, as we have definitely had success with the value proposition of research efficiency, but in general, a lot of things business adopt to improve project time (like Theory of Constraints project management, Mindset/Skillset/Toolset matching of personel et) is of no interest to academic researchers.

Re: A farewell to bioinformatics (2012)

#168
post #163

Earlier quoted context omitted.

Why wouldn't anyone buy your product? If it is easy to use, and SPEEDS UP RESEARCH TIME, your researcher/PI who is spending thousands on computing clusters will buy your software for their graduate students. Hell, my PI keeps asking me if I need a faster computer so I can run Matlab better/quicker. Really, if I had a software that helped me perform research faster/better/quicker and compare my results to ground truth…

Ahh the efficiency argument. The trick is, academics often have excess manpower capacity in the form of grad students and post-docs. Even though personell is usually one of the highest expenses on any given grant, they often don't look at ways to improve the efficiency of their research man-hours. That's not a blank rule, as we have definitely had success with the value proposition of research efficiency, but in gene…

After researching this field (biomedical R&D) a bit, I found that the mindset and workflow is mostly pre-computers. The relevant decision makers in the labs usually don't see a need to change something because "it works" and "it's done always this way".

Re: A farewell to bioinformatics (2012)

#169

Earlier quoted context omitted.

>There can be a big opportunity cost in trying to rework a workflow so that it is more efficient and then test it thoroughly ensure correctness. Hi, I recognize your name as a legit bioinformatician, am a huge fan of the lab that you're currently in, and others should listen to you. I'd like to add that for many projects, general reusable software engineering is not necessarily a huge advantage. Instead of verifying…

Software engineering is important for bioinformatics, in my opinion. But it's important to identify the things that are important and aren't: Reproducible code: extremely important. Correct code: extremely important. Readable code: very important. Efficient code: often not as important. Even today, the UCSC Genome Browser is an example where efficient code is important. It is interactive software, has many human user…

>Reproducible code: extremely important. Correct code: extremely important. Readable code: very important. Efficient code: often not as important.

You want Haskell. :)

Re: A farewell to bioinformatics (2012)

#170
The bio in bioinformatics is the important bit. Informatics plays second fiddle, even in the name. Very few will appreciate your beautiful code, but many will appreciate you finding a cure for cancer. That is the reality of bioinformatics, most of the code has a short shelf life. If you luck out, your software may live longer, as is the case with samtools. That samtools code is crappy is true, still the much cleaner code alternatives, sambamba and bamtools, are not much used! Go figure.

Maybe bioinformatics is not the place to aim for great informatics. We do bioinformatics because of love of science first and foremost. This is frontier land, the wild west, and it pays to play quick and dirty. I would suggest to hang on to some best practices, e.g. modularity, TDD and BDD, but forget about appreciation. Dirty Harry, as a bioinformatician you are on your own.

To be honest, in industry it is not much different. These days, coders are carpenters. If you really want to be a diva, learn to sing instead.

Post reply on HN