# Best practice for channels that use HDF5 (seems like race is still an issue with lock)

**URL:** <https://discourse.julialang.org/t/best-practice-for-channels-that-use-hdf5-seems-like-race-is-still-an-issue-with-lock/44652>\
**Category:** New to Julia\
**Tags:** jld, hdf5, multithreading, channel\
**Created:** [August 10, 2020, 1:07am UTC](https://discourse.julialang.org/t/best-practice-for-channels-that-use-hdf5-seems-like-race-is-still-an-issue-with-lock/44652 "2020-08-10T01:07:19Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkarikom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkarikom/32/7096_2.png) [@mkarikom](https://discourse.julialang.org/u/mkarikom)\
**Post date:** [August 10, 2020, 1:07am UTC](https://discourse.julialang.org/t/best-practice-for-channels-that-use-hdf5-seems-like-race-is-still-an-issue-with-lock/44652/1 "2020-08-10T01:07:19Z")

</div>

I’m about to submit some jobs that use following “emergency backup” scheme (see MWE testBackup.jl below).  
I learned the hard way after loosing some data on a job that was submitted to our cluster, due to a failing checkpoint server.

Per recommendation in the manual, [data-race freedom is provided](https://docs.julialang.org/en/v1/manual/multi-threading/#Data-race-freedom-1) by the function `runSimulation` which acquires a lock on the channel `chn1` controlling access to JLD.@save, but if I save too often (by lowering saveinterval), bad things happen during local testing like the repl crashing (see below), or rarely Nautilus crashing.

The `csize=10` argument on the constructor for `chn1` should be unnecessary, but I just put it in there to see if it helped make things more reliable (it didn’t).

Stuff that is getting backed up will typically be the serialized representation of the model state (very large), so JLD seemed like the way to go, vs say writing some values to a database. Is there a better way to do this?

Job finishes (saveinterval = 50):

 ![done](https://global.discourse-cdn.com/julialang/original/3X/b/d/bd88282b6f245d1a8c70de9760c8db1b5eadbf9c.png)

Job crashes (saveinterval = 5):  
 ![not](https://global.discourse-cdn.com/julialang/original/3X/b/f/bfb537819561d05f1e04529600455ed668cc01cc.png)

testBackup.jl:

```julia
using JLD

# hdf5 and therefore jld are not threadsafe, so run this to controll access to hdf5 library
function writeJLD(c::Channel)
  while true
    data = take!(c)
    JLD.@save data["fn"] data["state"]
  end
end

function runSimulation(a,b,c::Channel)
    for i in 1:1000
        # so stuff
        saveinterval = 5
        if mod(i,saveinterval) == 0
            fn = string(a,"_data.jld")
            print("\n saving to ",fn)
            lock(c)
            try
                put!(c,Dict("fn"=>fn,"state"=>b*2))
            finally
                unlock(c)
            end
        end
    end
end

chn1 = Channel(writeJLD;csize=10)
inits = rand(10)
for i in 1:length(inits)
    print(string("\n running foo ",i))
    if i < 0.5
        sleep(2)
    end
    Base.Threads.@spawn runSimulation(string("/tmp/foo",i),inits[i],chn1)
end

```
