Also see this comparison between Julia's BioSequences and Seq by Jakob Nissen and Ben Ward: https://biojulia.net/post/seq-lang/
Seq – A programming language for computational genomics and bioinformatics
21–30 of 59 posts
Re: Seq – A programming language for computational genomics and bioinformatics
#22I'm wondering if Seq can also serve as a general-purpose replacement for Python whenever a fast executable is needed.
Re: Seq – A programming language for computational genomics and bioinformatics
#23It’s odd that they didn’t include Nim in the benchmarks in their paper: https://dl.acm.org/doi/pdf/10.1145/3360551
I know nothing about Nim or genomics. Why is it odd that they didn’t include Nim?
Nim can be sold as a "A strongly-typed and statically-compiled high-performance Pythonic language" as Seq (although it is more than that and does not actually have as a goal to be Pythonic, see https://nim-lang.org/ or https://github.com/Araq/nimconf2021/blob/main/zennim.rst).
Still, given the small size of Nim community and even smaller size of the genomics nim subcommunity, I would say it is not that odd that is not included in the benchmark. The existing nim genomics library might not even cover the functionalities required by the benchmark.
Re: Seq – A programming language for computational genomics and bioinformatics
#24I really like that Seq seems to have built-in some parallelization ability. I spend no small amount of time in my day job doing that manually in R with RcppParallel for loops that are totally independent across each iteration.
Bioinformaticians are often educated to use a specific programming language and environment. They aren't usually looking to try other languages. For example, I support our bioinformatics group and they are basically 100% R and RStudio users. We have a single user of Python and that user is doing "typical" tensorflow stuff with images.
I've noticed this same bias towards a single language for some other academic niches. Like SAS or Stata camps in public health or psychology - I think of these languages as basically the same, but for non-CS folks the perception seems to be more like English vs Russian.
Even more complicated, researchers may be extremely committed to a specific library in a language and suspicious of languages that don't have their favorite library available.
Any shift to new tooling for these highly-committed users will almost certainly require large and obvious benefits to gain traction.
Re: Seq – A programming language for computational genomics and bioinformatics
#25I don't expect the community will adopt other languages at a large scale. My hope, though, is that more of these algorithms move to real distributed processing systems like Spark, to take advantage of all the great ideas in systems like that. But genomics will continue to trail the leading edge by about 20 years for the foreseeable future.
Re: Seq – A programming language for computational genomics and bioinformatics
#26Typically, any high performance (low latency or high throughput) genomics/bioinformatics applicaiton is not going to be written in plain Python, except possibly for prototyping. Instead, nearly all codes today are written in C++ or Java, with some sort of command and control in Python or a DAG-based workflow scheduler. I don't expect the community will adopt other languages at a large scale. My hope, though, is that…
Re: Seq – A programming language for computational genomics and bioinformatics
#27Typically, any high performance (low latency or high throughput) genomics/bioinformatics applicaiton is not going to be written in plain Python, except possibly for prototyping. Instead, nearly all codes today are written in C++ or Java, with some sort of command and control in Python or a DAG-based workflow scheduler. I don't expect the community will adopt other languages at a large scale. My hope, though, is that…
IMO, spark isn't the way forward. The typical pattern with it is it lets you scale up to 100 cores really easily which is almost enough to compete with a good single threaded implementation in a fast language.
The workflows I deal with generally involve moving hundreds of terabytes of storage into memory, processing it, and writing it out. Single machines (even beefy ones) tend to hit their limits (networking, max RAM, cache size, TLB, etc).
Maybe there's another tool better than spark, i don't know, the important thing is that spark is the most ubiquitous.
Re: Seq – A programming language for computational genomics and bioinformatics
#28Re: Seq – A programming language for computational genomics and bioinformatics
#29How do you pronounce Seq?