# DrWatson - the perfect sidekick to your scientific inquiries!

**URL:** <https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241>\
**Category:** Package Announcements\
**Created:** [April 17, 2019, 2:51pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241 "2019-04-17T14:51:02Z")\
**Posts on this page:** 16\
**Page:** 2

<div class="post-metadata">

**Author:** ![theogf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/theogf/32/1987_2.png) [@theogf](https://discourse.julialang.org/u/theogf)\
**Post date:** [April 18, 2019, 2:18pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/21 "2019-04-18T14:18:56Z")

</div>

This is is just amazing! I started to work on the exact same idea but ended up giving up!

An additional feature I had in mind was to have a GUI (based on Electron and Interact.jl) to prepare the simulation runs. This would include create/edit a template, and save/load specific config files.  
I would be happy to contribute in any case!

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [April 18, 2019, 2:22pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/22 "2019-04-18T14:22:06Z")

</div>

Great, good to have you on board!

I don’t fully understand your suggestion, so I think it is best to open a feature request issue to explain it in detail! I do want to comment though that Electron+Interact are quite heavy dependencies and also not only Julia dependencies which is something one should always think twice before adding. But of course it could be worth it.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [April 18, 2019, 8:31pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/23 "2019-04-18T20:31:46Z")

</div>

You tackle stashing files that don’t permit metadata in their format: Generate a filename that encodes parameter values. That kind of scheme has a second part: One needs to parse back the filename into its parameter values.

When I do this kind of thing in a project, then it is very annoying to generate a parameterset-\>filename mapping, plus regex to reconstruct the parameterset from the filename, in a way that is still human readable and does not lead to extremely long names. This is a giant ugly kludge, and kudos for trying to deal with it for us.

While I saw that you tackle name generation, I did not see any mention of parsing of names in the docs. Is that supported?

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [April 18, 2019, 9:13pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/24 "2019-04-18T21:13:07Z")

</div>

Thank you very much for your kind words. Parsing of names is not yet implemented, however I always thought about it. I believe this is a functionality we should have and that it is also easy to implement.

We just didn’t have the manpower to do it until this time. I’ve opened up an issue that summarizes the process: [https://github.com/JuliaDynamics/DrWatson.jl/issues/38](https://github.com/JuliaDynamics/DrWatson.jl/issues/38) contributions would be super welcome, otherwise I will do it as time permits! 🙂

---

<div class="post-metadata">

**Author:** ![grero](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/grero/32/109_2.png) [@grero](https://discourse.julialang.org/u/grero)\
**Post date:** [April 19, 2019, 12:10am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/25 "2019-04-19T00:10:52Z")

</div>

Interestingly , this is something I tried tackle, though inelegantly, in my [DataProcessingHierarchyTools](https://github.com/grero/DataProcessingHierarchyTools.jl) package. I basically needed a way to navigate a directory structure with the directory names defining a certain level of analysis. In my particular case, I am analysing neural data, and I have some analysis that run on an entire session, some on arrays of recording channels for that session, and some on individual cells. This tool allows me to automatically navigate to the appropriate level by defining a level parameter attached to each analysis type.  
What I ended up doing for parameters was simply to attach a hash of those parameters to the file name, so that when I run analysis with identical arguments, the results are simply loaded. Of course, this means that I can’t tell what the arguments were simply by looking at the filename. Anyway, DrWatson seems to be much more polished version of this, and as I said before, I’ll try to integrate that into my workflow.

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [April 19, 2019, 8:14am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/26 "2019-04-19T08:14:23Z")

</div>

> [@ExpandingMan](#):
>
> By the way, a “new variant” of this problem for me is that now I frequently have to save to S3 buckets rather than the filesystem (whether actual or emulated).

[DVC](https://dvc.org/) works quite nicely for me for versioning large files using S3 as the storage. It’s clearly focused on predictive modelling workflows but the versioning system is generic so it could fit different use cases.

---

<div class="post-metadata">

**Author:** ![theogf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/theogf/32/1987_2.png) [@theogf](https://discourse.julialang.org/u/theogf)\
**Post date:** [April 23, 2019, 11:15am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/27 "2019-04-23T11:15:10Z")

</div>

This is what I had in mind (this is WIP)  
That would be the template creator. Then one could just select parameters and save them in a config file.

 ![demo_config_file](https://global.discourse-cdn.com/julialang/original/3X/8/4/8454d7df175b625dad8760f5ecb3099bb5e43bed.gif)

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [April 23, 2019, 11:24am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/28 "2019-04-23T11:24:01Z")

</div>

Hey, this seems cool. But can you explain its purpose? What is this GUI supposed to achieve? (i.e. what does one do with the saved config file?)

Also, what are all these fields, like field name? I can see that the field name you wrote has space so it can’t be a Julia variable.

---

<div class="post-metadata">

**Author:** ![theogf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/theogf/32/1987_2.png) [@theogf](https://discourse.julialang.org/u/theogf)\
**Post date:** [April 23, 2019, 11:29am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/29 "2019-04-23T11:29:47Z")

</div>

Well the idea would be to first have a template config file for a project.  
Then one could create specific config files for each experiment that would directly be fed to the workspace to run an experiment, similarly to your `dict_list` function.  
From my experience I find it easier and less error-prone to work in a GUI to select parameters instead of manipulating dictionaries directly. It also allows to have the config file saved in the results folder as well.

And the field are not directly julia variables. They would simply be fields names.

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [April 23, 2019, 11:41am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/30 "2019-04-23T11:41:18Z")

</div>

I see, but to _really_ understand this, I would still need a usage demonstration or at least explanation. What do you do with the config file? How do you use it? What is the config file? is it XML, Julia, Toml? What’s its type? How do you actually use it in a simulation? Also, don’t you need to write a special parser for this to work?

For the dictionary all these questions are immediately answered since it is a basic Julia structure.

> [@theogf](#):
>
> From my experience I find it easier and less error-prone to work in a GUI

Yes this is a valid point, but one should consider that you may need to do these things over a cluster, or a cloud, or any other connection that won’t be able to support this. This is an advantage of the dictionary approach. A second advantage is that it works consistently with any conceivable type, existing or not (due to how we handle `Vector` subtypes). A final point is simply that Electron is a very heavy dependency.

* * *

Please notice: I am not bashing you or anything. From personal experience, the best way to improve something is to be as critical as you can, which is what I do here.

Do you have the code for this somewhere?

---

<div class="post-metadata">

**Author:** ![theogf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/theogf/32/1987_2.png) [@theogf](https://discourse.julialang.org/u/theogf)\
**Post date:** [April 23, 2019, 1:41pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/31 "2019-04-23T13:41:51Z")

</div>

Thanks for the helpful feedback. I did not think the whole thing through but my pipeline idea would be :  
Create Template Config File → Save as JSON (contains field name, default value and limits/options)  
Create Config File → Open existing Template file → Set values → Save as a `Dict` (where field names are keys) in JSON.  
The config file can then be directly fed by being read as a `Dict`.  
In short it is simply a practical (?) config file GUI maker 😅

I advanced a bit more to make things maybe more clear

 ![full_demo](https://global.discourse-cdn.com/julialang/original/3X/3/3/33b732953d3c9124862b758398290dfd82c90231.gif)

For the code it’s pretty ugly so far but you can still check it out if you want :  
[https://github.com/theogf/MLExps.jl](https://github.com/theogf/MLExps.jl)  
As you see I was going in the same direction as you did. You can simply check `src/gui_config.jl` and `test/test_gui.jl`

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [April 23, 2019, 2:06pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/32 "2019-04-23T14:06:02Z")

</div>

Thanks for the responce! To keep this post as on-topic as possible, I’ve continued further points in the repo you shared: [https://github.com/theogf/MLExps.jl/issues/1](https://github.com/theogf/MLExps.jl/issues/1)

---

<div class="post-metadata">

**Author:** ![wlandau](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wlandau/32/8993_2.png) [@wlandau](https://discourse.julialang.org/u/wlandau)\
**Post date:** [June 21, 2019, 11:25am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/33 "2019-06-21T11:25:56Z")

</div>

Author of `drake` chiming in here. TL;DR: I think the most direct apples-to-apples comparisons here are DrWatson vs MLFlow and `drake` vs [GNU Make](https://kbroman.org/minimal_make/).

> Is this something similar to [https://ropensci.github.io/drake/](https://ropensci.github.io/drake/)

DrWatson has different goals: version control, sharing, project file structure, and provenance (keeping track of simulation settings). In these respects, it is more like MLFlow, especially when it comes to tracking. `drake`, on the other hand, tries to be Make for R. `drake` synchronizes expensive computations in an end-to-end pipeline so repeated full runs take minimal time. `drake` analyses your targets and functions to figure out what needs to run and what can be skipped, taking into account that some computational steps depend on others,.

 ![Screenshot_20190621_072130](https://global.discourse-cdn.com/julialang/original/3X/c/d/cd2268aa72409182ada3d0b443e9ad5922417304.png)

I will say that DrWatson and `drake` are similar in that they both abstract away output file management and reduce manual bookkeeping.

> No, from the 6-minute video I just watched they don’t seem similar. DrWatson also seems to have a much simpler and cleaner approach to helping you (e.g. you don’t have to make a “DataFrame” out of every function in your code!)

I disagree with that characterization of the data frame (the `drake` plan). You do not need to wrap up all your functions in it. In `drake`, you can define your supporting functions wherever and however you want, and then your commands in the plan simply reference those functions as needed. Most of your code lives in the functions, as is the case for cleanly-implemented scientific workflows in general, `drake` or no `drake`.

The `drake` plan is like a `Makefile` for R. The main differences are

1. You are writing R code.
2. The syntax is _much_ easier than Make wildcards.
3. You do not need to list out all the dependencies of each target manually by hand. `drake` automatically discovers dependency relationships among your targets and functions using static code analysis. (See the graph above.)

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [June 21, 2019, 11:37am UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/34 "2019-06-21T11:37:09Z")

</div>

> [@wlandau](#):
>
> The `drake` plan is like a `Makefile` for R. The main differences are

I’m not familiar with `drake`, but is it analogous to snakemake? There’s also Makeitso.jl… wondering where that fits in.

---

<div class="post-metadata">

**Author:** ![wlandau](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wlandau/32/8993_2.png) [@wlandau](https://discourse.julialang.org/u/wlandau)\
**Post date:** [June 21, 2019, 12:39pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/35 "2019-06-21T12:39:11Z")

</div>

Yes, `drake` is much more similar to `snakemake` and `Makeitso.jl`.

---

<div class="post-metadata">

**Author:** ![Datseris](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/datseris/32/13406_2.png) [@Datseris](https://discourse.julialang.org/u/Datseris)\
**Post date:** [June 23, 2019, 3:22pm UTC](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241/36 "2019-06-23T15:22:47Z")

</div>

Hi, @wlandau, thanks for chiming in. Welcome to the Julia discourse! Glad that this post somehow made its way to you so we can get a more fair representation of the other software mentioned here!

> [@wlandau](#):
>
> I disagree with that characterization of the data frame (the `drake` plan).

Okay, I can accept this. I didn’t spend a lot of time to learn `drake` so I may not be accurate in my judgement. But be aware though, that the workflow dependency graph that you showcase in your original comment is already too complicated for me as a working scientist and also compared to what DrWatson needs to achieve. I think we just have different goals and/or target groups.

> [@wlandau](#):
>
> DrWatson has different goals: version control, sharing, project file structure, and provenance (keeping track of simulation settings).

Correct but, lets not forget the “Naming Simulations” part (see the [Functionality](https://juliadynamics.github.io/DrWatson.jl/dev/#Functionality-1) page), which is actually what I personally use most often from DrWatson!

[Previous page](https://discourse.julialang.org/t/drwatson-the-perfect-sidekick-to-your-scientific-inquiries/23241.md?page=1)
