# Snakemake with julia (and similarities with Dr Watson or..)

**URL:** <https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106>\
**Category:** General Usage\
**Tags:** workflow, scientific-project, drwatson, reproducibility\
**Created:** [February 20, 2025, 9:36am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106 "2025-02-20T09:36:29Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Post date:** [February 20, 2025, 9:36am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/1 "2025-02-20T09:36:29Z")

</div>

I have a collegue that works with Python and [Snakemake](https://snakemake.github.io), telling me how nice is this tool for scientific workflow management, but I haven’t understood much about it, and above all if what it provides is really needed much more in a python world rather than in a Julia one (for example, reproducibility/containerisation).

In the website of Snakemake they cite also Julia, is there anyone that use it with Julia? For which reasons? Do you have a public example to share ?

Is [DrWatson](https://juliadynamics.github.io/DrWatson.jl) similar in objectives but more tailored to Julia workflows, or is really something different ?

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [February 20, 2025, 10:40am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/2 "2025-02-20T10:40:44Z")

</div>

In the [DrWatson paper](https://joss.theoj.org/papers/10.21105/joss.02673) we do discuss alternatives in other languages, but unfortunately snakemake isn’t there. Maybe it is a recent development? Would love to hear a summary from someone that have used it.

---

<div class="post-metadata">

**Author:** ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Post date:** [February 20, 2025, 10:43am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/3 "2025-02-20T10:43:22Z")

</div>

Not so new…

 ![image](https://global.discourse-cdn.com/julialang/original/3X/b/1/b11450656a5fae7af6ef56647aec16e26fd77173.png)

I would love too 🙂🙂🙂

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [February 20, 2025, 10:59am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/4 "2025-02-20T10:59:04Z")

</div>

I was checking it out now. The framework focuses on reproducible (and cross-platform?) data analysis pipelines. DrWatson isn’t really about data management / analysis which may be the reason we didn’t compare in the paper? I don’t remember it was many years ago!

---

<div class="post-metadata">

**Author:** ![mrufsvold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mrufsvold/32/31600_2.png) [@mrufsvold](https://discourse.julialang.org/u/mrufsvold)\
**Post date:** [February 20, 2025, 11:29am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/5 "2025-02-20T11:29:48Z")

</div>

I believe @tecosaur 's DataToolkit.jl would be the Julia competitor in this space!

---

<div class="post-metadata">

**Author:** ![abraemer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abraemer/32/51403_2.png) [@abraemer](https://discourse.julialang.org/u/abraemer)\
**Post date:** [February 20, 2025, 6:28pm UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/6 "2025-02-20T18:28:33Z")

</div>

Funnily enough there was a post 2h before this thread about basically the same topic by @jonathanBieler

> [@Dagger + Dates = snakemake?](https://discourse.julialang.org/t/dagger-dates-snakemake/126111):
>
> I have a folder data with samples (csv files), I want to compute some statistics for each file and then make a summary. In snakemake you can define rules how to produce the output file from the inputs, and snakemake will handle the execution, running things only when needed (e.g. when a file has been modified). There’s not equivalent in Julia, but I think all most of the building blocks are already in Dagger and we just need some nicer frontend. Here’s an example I made quickly to illustrate: I…

---

<div class="post-metadata">

**Author:** ![jonathanBieler](https://avatars.discourse-cdn.com/v4/letter/j/82dd89/32.png) [@jonathanBieler](https://discourse.julialang.org/u/jonathanBieler)\
**Post date:** [February 20, 2025, 8:58pm UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/7 "2025-02-20T20:58:25Z")

</div>

As I understand DrWatson is a set of tools for setting up and easily do some commons operations in a reproducible way, but it’s not a workflow system. In snakemake, nextflow, etc or the dagger example I posted, one of the goal is to build a graph that capture the dependencies between the inputs files and the results you want to compute, and to execute that workflow automatically in parallel, with the ability to resume, or update only the part that needs to be updated.

e.g. if you have a workflow that starts from data A, compute B from it, and then make a plot C, you can represent it like this :

A → B → C

In Julia you could make a script like this :

```julia
include("make_A.jl")
include("make_B.jl")
include("make_C.jl")

```

Now let’s say you want to modify B but you don’t want to recompute A (it takes two days) so you modify your script :

```julia
#include("make_A.jl")
include("make_B.jl")
include("make_C.jl")

```

Two weeks later the data A have been updated so you rerun your script but you forgot that you commented the first script, so you silently don’t get the results you intended.

Of course that’s a simple example, but I’m sure many people have experienced these kind of problems in practice. DataToolkit seems to be doing some of that, but it seems more tailored to managing/ingesting the raw data that being a full workflow system with parallel execution, monitoring, etc.

The issue with snakemake or nextflow is that they add quite a bit of overhead, they have their own way of doing things you have to learn and adapt to, which are very useful in some contexts, but for “non-industrial” data science I would prefer something more lightweight and flexible with the minimum amount of boilerplate possible.

---

<div class="post-metadata">

**Author:** ![tecosaur](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tecosaur/32/23206_2.png) [@tecosaur](https://discourse.julialang.org/u/tecosaur)\
**Post date:** [February 21, 2025, 2:07am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/8 "2025-02-21T02:07:00Z")

</div>

> [@mrufsvold](#):
>
> I believe @tecosaur 's [DataToolkit.jl](https://juliaregistries.github.io/General/packages/redirect_to_repo/DataToolkit) would be the Julia competitor in this space!

Thanks for the shoutout! 😍

> [@jonathanBieler](#):
>
> DataToolkit seems to be doing some of that, but it seems more tailored to managing/ingesting the raw data that being a full workflow system with parallel execution, monitoring, etc.

Yup. The way I’d put it is that DataToolkit has some of the _pieces_ of a workflow system, but it isn’t one. There’s a similar story with Dagger.

That said, I designed DataToolkit to be able to do things I didn’t plan for it to be able to do, and I think it’s very much possible for it to gain some of the key missing pieces of functionality.

For example, I’ve been thinking of making a plugin that allows for “parametric data sets”, and I also have a few thoughts on how to make it so it can run across multiple machines at once.

Oh, to give an example of what it can currently do, I might as well show off the MetaGraphsNext extension:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/a/2/a27937eb89707f04482dd7b8be09e73a74939906.png)

This shows all the datasets I used for a project, as well as the dependencies between them.

There’s potential for a plugin that integrates with some parallelisation tool to split up dependencies and execute them in parallel, but nothing currently developed.

---

<div class="post-metadata">

**Author:** ![tecosaur](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tecosaur/32/23206_2.png) [@tecosaur](https://discourse.julialang.org/u/tecosaur)\
**Post date:** [February 21, 2025, 5:24am UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/9 "2025-02-21T05:24:03Z")

</div>

> [@jonathanBieler](#):
>
> one of the goal is to build a graph that capture the dependencies between the inputs files and the results you want to compute, and to execute that workflow automatically in parallel, with the ability to resume, or update only the part that needs to be updated.

Just for clarity, what DataToolkit currently supports OOTB is:

- build a graph that capture the dependencies between the inputs and the results
- parallel workflow execution
- the ability to resume a workflow
- incremental/minimal updates
- generic processing steps, that can be applied in bulk

If anybody is interested in helping tick a few more boxes, I’d be very happy to collaborate.

---

<div class="post-metadata">

**Author:** ![moble](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moble/32/23535_2.png) [@moble](https://discourse.julialang.org/u/moble)\
**Post date:** [February 23, 2025, 8:39pm UTC](https://discourse.julialang.org/t/snakemake-with-julia-and-similarities-with-dr-watson-or/126106/10 "2025-02-23T20:39:57Z")

</div>

I think everyone interested in any of this might also be interested in [showyourwork](https://show-your.work/en/latest/), which _uses_ snakemake, but does a lot more.

@MilesCranmer had a very nice series of posts starting [here](https://discourse.julialang.org/t/is-there-a-julia-package-similar-to-the-pythons-showyourwork/97475/7) discussing this, comparing to other tools (including the excellent DrWatson, which I particularly love!), and even providing an example for use with Julia.
