Live data from Hacker News

Consider working on genomics

claymcleod.dev

191–200 of 299 posts

Re: Consider working on genomics

#191
post #72

20 years ago I got interested in "bioinformatics." I loved learning something about molecular biology, after all those years of hearing about DNA and not understanding it. And "Molecular Biology of the Cell" is, hands down, the greatest textbook ever written. That said: a lot of the comments are spot on. You're working in a field where the hard scientists and business people rule and you're a helper. Maybe they're gr…

>Molecular Biology of the Cell Tangential, but what are the chemistry prereqs to grasp this book?

Probably just a college level gen chem class. Pretty accessible, albeit technical, textbook from what I remember of reading it for a course a few years ago.

Re: Consider working on genomics

#192

Quick plug here for Atomic AI ( https://atomic.ai/ , https://boards.greenhouse.io/atomai ), which could be added to the list. We value and respect (and pay) our engineers—I myself trained as a SWE and worked at FAANG. Shoot me a message at raphael@atomic.ai if you want to learn more.

(chiming in here as a founding engineer at Atomic)

So I spent more than 8 years as a SWE at Google, and now work here with both experimental biologists and machine learning scientists. And yes, a lot of the concerns mentioned in this thread are also things I have had anxiety about.

Most obvious to me, being a software engineer at Google felt like being the center of the universe. Coming here, the focus is the scientific research. And yes, the scientists all managed to complete their PhDs so they don't necessarily need me to unblock them every second of their day. But contrary to my expectations, this has been remarkably freeing. I think one particularly important part of our company that makes this work is that, even on the science side, we're multidisciplinary (at a high level, emphasizing both experimental biology and ML.) And so engineering feeling like another arm of that multi-discipline nature is fairly... natural.

The reason I feel it's freeing, and the reason I enjoy working here, is also the greatest challenge. Because the scientists are focused on the science, because they respect me and trust me to figure it out, and because they aren't constantly blocked by me, my job is mostly about dreaming extremely expansively about what I can do to reduce toil and make the scientists more productive. Of course they have feedback and input, but how I use my time and what I build is ultimately my decision because I am the engineer. And I have been able to do some things I am very proud of, like rolling out Bazel and Kubernetes and finding ways to seamlessly bring them into the cloud (we're even multi-cloud now without them even noticing!) On the other hand, it's very challenging because when you work on a product, say Google Photos, as a SWE, you always have some direct tether to the product ("what should we build next? ahhhh, well I guess we could just embed stable difficusion and a million people would immediately play with it".) At Atomic, my tether is very ambiguous. If I do my job successfully, they'll be able to do research more quickly (? effectively?), and eventually we'll be able to produce a therapeutic that hopefully changes the world. Identifying what I can do today to speed up that far outcome in the future is very challenging, but it is a far more interesting challenge than gluing some pre-existing software into my UI or running A/B tests to turn a red button blue.

If, like me, you enjoy being given ownership over incredibly ambiguous problems, please do reach out!

This role focuses on directly partnering with the biologists: https://boards.greenhouse.io/atomai/jobs/4726839004

This role is expansive cloud infra: https://boards.greenhouse.io/atomai/jobs/4531035004

And this role is directly partnering with the ML scientists: https://boards.greenhouse.io/atomai/jobs/4191285004

Re: Consider working on genomics

#193
Not quite a decade ago, I took some work for a lab to replace some aging software (circa 1990) used to do peptide synthesis.

It was an enlightening experience. While I was the programming expert with a CS degree, I wasn't trusted for anything, because I wasn't a PhD or had a background in bioinformatics. However, I did get to work with lots of smart people, fixed and improved the code and processes that the Phd level statisticians and bioinformaticians used.

It is a real joy to work in hard science, with brilliant people who love their work. I learned a ton and gained a healthy respect for the people that do this kind of work.

However, the downsides are pretty bad. Pay and compensation is awful. Most people, myself included, could have made as good if not better pay waiting tables. There end up being different levels of people Administrators, Private investigators, and lab workers (peons). Unless you are an admin or a high level PI you're not gonna be getting much money.

Everybody lives and dies by the grant. If funding dries up, you will be out of a job.

Ethics. Us CS people are woefully under educated on ethics. You will find yourself asking why we can simply do something, often the answer will be ethics.

Regulations, like ethics, you will have to bend to regulations. It's not a bad thing, just a different thing.

Unless you find yourself in a admin role, you will just be another lab peon. Its not a a bad place to be, but you will never be at the top of the totempole.

Loads and loads of ego. You will work with very smart and sometimes unreasonable people. Learning to navigate this with tact is important.

Re: Consider working on genomics

#196
As someone who puts tremendous value in technical mentorship when considering a role this is about the worst possible advertisement for being a swe in genomics as it amounts to "all our code is awful- come fix it!"

Re: Consider working on genomics

#197

"From my experience, what works incredibly well is a partnership between biologists and software engineers: the biologists first come up with the first concept of the tool, which is purely focused on ensuring good results. After this first iteration is completed, engineers then come in and rewrite the tool using modern engineering practices with things like speed and reliability in mind." Like others have pointed out…

Yeah I think this is fair enough after reading it back. However, that was not exactly my intention here, and I think this is a case of me needing to be more careful in my wording.

When I said that software engineers add in the speed and reliability, I didn't mean they _only_ add in the speed and reliability: just that these two tenants of good software engineering where accounted for in this "correct" way of doing things (as opposed to the state of most genomics software that I described above).

However, I can see how my phrasing can give the wrong impression about the contributions an engineer makes when the biologist and engineer sit down to do create the real thing together. In a positive environment, both sides (biologists and software engineers) share enough information with one another that the either can make contributions to the scientific/software engineering domain.

Re: Consider working on genomics

#198

Can you provide a list of the top problems in that space? Much rather try to understand them deeply myself and build a company solving them than just getting a job.

Protein structure prediction was a huge deal, which is why AlphaFold received so much fanfare. It is actually pretty good. The next step is to predict where multi-protein complexes would interact- which is not just as simple as predicting the structure of two proteins independently and then trying to fit them together like a puzzle, because the the interactions can also change the structure. While it's not as hard as it used to be to experimentally determine protein targets of, for example, a protein kinase, it's still not an arbitrary or cheap experiment, and to do that for the many thousands of such proteins, across different conditions (stress, presence of co-factors, etc) and in different organisms would be rather a lot of work. Something like alphafold that makes reasonable predictions and can be used to help you focus on what's most likely to be relevant to your disease or process of interest helps quite a bit.

There's also more need for integrating "multi-omics" data, where you have data from multiple assays (gene expression, phospho-proteomics, lipidomics, epigenetics, small RNA expression, etc etc) with the goal of somehow combining all these different assay results from various levels of gene regulation, to get closer to figuring out actual mechanism for complex processes. Building on that, we can also do single-cell multi-omics to some extent- where you have results from different sequencing-based assays on the level of the same individual cell. This is still pretty limited, but it's exciting and advancing pretty quickly. This will eventually be combined with things like spatial transcriptomics, which is useful for mapping out what's going on in heterogeneous tissue samples like tumors, for example, so we'll end up with spatial single-cell multi-omics, at which point you're looking at 1) some quantitative trait for multiple genes/loci/molecules, and often 10k+ of such features at the same time per assay, 2) multiple assays, such as DNA accessibility and gene expression, in 3) single-cells, of which you might have 10k of in a single sample, 4) across a physical tissue sample where individual cells are spatially mapped, and where you probably want to figure out how cells might influence the state of those around them, and 5) in multiple different samples, where you might want to compare disease vs control, or look for correlation to heterogeneity of results within one group.

There's a lot of public data already available for single-cell gene expression projects if you want to get a feel for how these things are structured and how (passable but not amazing) the existing tooling is- one of the main repositories for this data is the NCBI's SRA https://www.ncbi.nlm.nih.gov/sra but you'll quickly note that searching and browsing is not as easy as you might think it would be- because one of the main limiting factors in bioinformatics is how bad everyone is at keeping terminology consistent. For many bioinformaticians, a majority of time is spent in the data cleaning phase. It's awful. Sometimes the experimental parameters make it into SRA or GEO, but sometimes you have to read through the associated paper to pull that out. Often it's only large consortium projects like the The Cancer Genome Atlas (TCGA) or the Genotype-Tissue Expression project (GTEx) - which have enough funding for staff dedicated to data management- end up publishing datasets that are easy to "consume" without having to jump through a whole bunch of hurdles to figure out how the data was produced.

I have a BS/MS in bioinformatics and I'm presently a PhD candidate in genetics and computational biology defending in February.

Re: Consider working on genomics

#199

Google already axed all job offers, Microsoft and AWS are searching student interns... I used to work in genomics and computational biology. It was incredibly interesting. But it's university research and gets paid as such. 2-year time-limited contracts, lots of interns and students, extremely low salaries.

The AWS jobs aren’t even related to genomics. They just have genomics in the description of types of workloads performed by customers of AWS. The jobs are hard core CS automated reasoning jobs.

Re: Consider working on genomics

#200

So if one was financially independent and wished to write something open-source in that field, where would the highest impact be?

Invent a new file format (or a few) for storing genomics data. They're all the rage in the bioinformatics field. Make sure not to document its semantics so that its implementation is the only spec.
Post reply on HN