Live data from Hacker News

Research software code is likely to remain a tangled mess

shape-of-code.coding-guidelines.com

101–110 of 171 posts

Re: Research software code is likely to remain a tangled mess

#101
post #59

Earlier quoted context omitted.

As someone who has been on both the research and industry software end, there’s really not that much difference. Requirements change, you build that into your plans. Frankly, a lot of best practice software development that gets totally ignored by academia (e.g. OOP) can handle this exact case, and makes things way more flexible. If the problem was only unpredictability, then projects with a clear and defined end goa…

I can tell you why the sites went offline, because the funding stopped. I don't know what you're research background is but its painful to even get 5 GBP a month to host a droplet on digital ocean in a pretty lucrative department with liberal internal funding.

Agreed, but all these little things are just a sign that the industry just does not give a shit about software. They could develop mechanisms to fund this stuff, pretty easily actually. But they don’t.

A couple of other weird inequities that I’ve found are: 1. It’s hard to get permission to spend money on software subscription based licenses since you won’t “have anything” at the end. However, it’s much easier to get funding for hardware with time based locks (e.g after 3 years the system will lock up and you have to pay them to unlock). The end result is the same, you can’t use the hardware after the time period is up, but for some reason the admin feels much more comfortable about it.

2. It’s hard get funding to hire someone to set up a service to transfer large amounts of data from different places. It’s much easier to hire someone to drive out to a bunch of places with a stack of hard drives and manually load the data on them, and drive back. Even if it’s 2x more expensive and would take longer. Why? Again my speculation is that the higher ups are just more comfortable with the latter strategy. They can picture the work being done in their head, so they know what they’re paying for.

Re: Research software code is likely to remain a tangled mess

#102
Keep in mind that there're different kinds of research software. Take Seurat[1] as an example. There's CI, issue tracking, etc. It might not be the prettiest code you ever seen, but it absolutely has to be maintainable as it's being actively developed. Such projects are rare, but the low quality is often an indication of a software that isn't used by anyone.

1. https://github.com/satijalab/seurat

Re: Research software code is likely to remain a tangled mess

#103
post #52

Earlier quoted context omitted.

> At the same time, you need to consider that such a clean up is only realistically helpful for other people to check whether there are bugs in the original results, and not much else. I assumed that most code published could be directly useful as an application or a library. Considering what you're saying, this might be only a minority of the code. In that case, I agree with your conclusion about smaller gains.

Yeah my standard quote about research code is that it is not the product, so it is ok thta it is bad. The results are the product and those need to be good. Someday someone will take those results (in the form of some data or a paper) and make a software product, and that should be good.

[deleted]

Re: Research software code is likely to remain a tangled mess

#104
One of the biggest eye-openers for me as an undergrad was when, upon getting to the point where I'd have to decide whether to pursue graduate education or exit academia and join the workforce, I began to look at the process for publishing novel computer science.

To be clear, novel computer science is valuable and the lifeblood of the software engineering industries. But the actual product? I discovered of myself that I like quality code more than I like novel discovery, and the output of the academic world ain't it. Examples I saw were damn near pessimized... not just a lack of comments, but single-letter variables (attempting to represent the Greek letters in the underlying mathematical formulae) and five-letter abbreviated function names.

I walked away and never looked back.

If there's one thing I wish I could have told freshman-year me, it's that software as a discipline is extremely wide. If you find yourself hating it and you're surprised you're hating it, you may just be doing the kind that doesn't mesh with your interests.

Re: Research software code is likely to remain a tangled mess

#105
post #15

If you want to write good research software, a good way is to have professional developers implement it. I worked closely with an NLP researcher for a while on a project that had received a hefty state grant. She knew more or less what her team needed, but she needed someone to implement it cleanly and in a way that would not make users step on each others toes. The chances of that project being a buggy mess would ha…

Here's the problem with hiring a pro.

The workhorse NIH grant[0] is a R01 with a $250,000/year x 5 years "modular" budget. Most labs have, at most, one. Some have two, and a very few have more than that. This covers everything involved in the research: salaries (including the prof's), supplies, publication fees, etc. Suppose you find a programmer for $75k. With benefits/fringe (~31% for us, all-in), that's nearly $100k/year. If the principal investigator (prof, usually) takes a similar amount out of the grant, there's very little money left to do the (often very expensive) work. In contrast, you can get a student or postdoc for far less--and they might even be eligible for a training grant slot, TAship, or their own fellowship, making their net cost to the lab ~$0.

This would be easy to fix: the NIH already has a program for staff scientists, the R50. However, they fund like two dozen per year; that number should be way higher.

[0] Other mechanisms exist at the NIH--and elsewhere--but NSF (etc) grants are often much smaller.

Re: Research software code is likely to remain a tangled mess

#106
post #38

There's a huge digital divide forming as well. Between the hardware a junior software engineer at a well funded research institution such as DeepMind has access to. Compared to the postdoc in Theoretical Physics at Princeton. Who is expected not only to write software. But maintain hardware for a proprietary "supercomputer" that was probably cast off ages ago from a government lab or wall street. We don't expect Aero…

> We don't expect Aerospace / Mechanical engineering students to learn metalworking.

Umm...we sorta do.

As a neuroscience postdoc, I have done virtually everything from analysis to zookeeping, including some (light) fabrication. We outsource really difficult or mass-production stuff to pro, and there's a single, very overworked machinist who can sometimes help you, but most of the time it's DIY.

Re: Research software code is likely to remain a tangled mess

#107
post #59

Earlier quoted context omitted.

I can tell you why the sites went offline, because the funding stopped. I don't know what you're research background is but its painful to even get 5 GBP a month to host a droplet on digital ocean in a pretty lucrative department with liberal internal funding.

Agreed, but all these little things are just a sign that the industry just does not give a shit about software. They could develop mechanisms to fund this stuff, pretty easily actually. But they don’t. A couple of other weird inequities that I’ve found are: 1. It’s hard to get permission to spend money on software subscription based licenses since you won’t “have anything” at the end. However, it’s much easier to get…

> The end result is the same, you can’t use the hardware after the time period is up, but for some reason the admin feels much more comfortable about it.

Simple: predictability. With a subscription based model, admin has to deal with recurring (monthly / yearly) payments, and the possibility is always there that whatever SaaS you choose it gets bought up and discontinued. Something you own and host yourself, even if it gets useless after three years, does not incur any administrative overhead and there is no risk of the provider vanishing. Also, there are no "surprise auto renewals" or random price hikes.

> 2. It’s hard get funding to hire someone to set up a service to transfer large amounts of data from different places.

Never underestimate the bandwidth of a 40 ton truck filled with SD cards. Joke aside: especially off-campus buildings have ... less than optimal Internet / fibre connections and those that do exist are often enough at enough load to make it unwise to shuffle large amounts of data through them without disrupting ongoing operations.

Re: Research software code is likely to remain a tangled mess

#108
post #98

Well as somebody who has written research software, I don't agree that research software is a "tangled mess". A couple of points, 1. often when I read read software written by profession programmers I find it very hard to read because it is too abstract, almost every time I try to figure out how something works, it turns out I need to learn a new framework and api, by contrast research code tends to be very self cont…

As someone who's worked for a large part of my career as a sort of bridge between academia and industry (working with researchers to implement algorithms in production), both you and the original author are right to an extent. On one hand, academics I've worked with absolutely undervalue good software engineering practices and the value of experience. They tend to come at professional code from the perspective of "I'…

> a lot of the smartest software engineers I've known have a terrible tendency to over-engineer things.

Your definition of "smartest software engineers" is the opposite of mine. In my view, over-engineering is the symptom of dumb programmers. The best programmers simplify complex problems; they don't complicate simple problems.

Re: Research software code is likely to remain a tangled mess

#109

I'm currently refactoring a fairly large piece of research code myself. It was written with lean startup thinking in that a little code ought to produce some value in its results. If i was able to eeek some usefulness out of this code, then Id put more energy into it. Otherwise I was perfectly happy to Fail Fast and Fail Cheap. How did it become such a mess in the first place? Simple - I didn't know my requirements w…

I like Brooks' "plan to throw one away; you will, anyhow.":

This [first] system acts as a "pilot plan" that reveals techniques that will subsequently cause a complete redesign of the system.

However, in practice I'm not confident enough in my understanding, and fear losing all that hard-won work, so I refactor too.

A rewrite from scratch is probably more viable when the project is small enough to keep in your head at once.

Re: Research software code is likely to remain a tangled mess

#110

Well as somebody who has written research software, I don't agree that research software is a "tangled mess". A couple of points, 1. often when I read read software written by profession programmers I find it very hard to read because it is too abstract, almost every time I try to figure out how something works, it turns out I need to learn a new framework and api, by contrast research code tends to be very self cont…

I work on mathematical modeling, dealing with human physiology. Likewise, the software packages used can be esoteric, and the structure of your "code" can be very different looking, to say the least.

This is certainly a lot of work, and this takes a lot of practice to perform efficiently: But no matter what, I comment every single line of code, no matter how mundane it is. I also cite my sources in the commenting itself, and I also have a bibliography at the bottom of my code.

I organize my code in general with sections and chapters, like a book. I always give an overview for each section and chapter. I make sure that my commenting makes sense for a novice reading them, from line-to-line.

I do not know why I do this. I guess it makes me feel like my code is more meaningful. Of course it makes it easier to come back to things and to reuse old code. I also want people to follow my thought process. But, ultimately, I guess I want people to learn how to do what I have done.

Post reply on HN