# Looping over a thousand files and checking if empty

**URL:** <https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291>\
**Category:** Performance\
**Tags:** multithreading\
**Created:** [December 7, 2023, 4:40pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291 "2023-12-07T16:40:04Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 4:40pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/1 "2023-12-07T16:40:04Z")

</div>

Hello!

I have some results of an experiment. They are stored inside a folder and I need to open the files, check what is the fraction of missing values inside a file and tag it according to different level of missing values (NaNs values actually). I’d like to make it parallel or at least speed up the code but so far I have not succeeded, any idea would be welcomed!

The number of file is about 2000 files where each file has a size of 16GB.

Note that I use a cluster, so if I can parrallelize, I could use a good number of CPU (using a Slurm scheduler).

Here’s a MWE (I hope!).

Project.toml content.

```julia
[179af706] CFTime v0.1.2
[336ed68f] CSV v0.10.11
[13f3f980] CairoMakie v0.10.8
[35d6a980] ColorSchemes v3.23.0
[3da002f7] ColorTypes v0.11.4
[5ae59095] Colors v0.12.10
[0703355e] DimensionalData v0.24.13
[61d90e0f] GeoJSON v0.7.2
[db073c08] GeoMakie v0.5.1
[c27321d9] Glob v1.3.1
[ee78f7c6] Makie v0.19.8
[85f8d34a] NCDatasets v0.12.17
[30363a11] NetCDF v0.11.7
[91a5bcdd] Plots v1.39.0
[438e738f] PyCall v1.96.1
[ac1d9e8a] ThreadsX v0.1.11
[5c2747f8] URIs v1.5.0
[90b8fcef] YAXArrayBase v0.6.1
[c21b50f5] YAXArrays v0.5.0
[0a941bbe] Zarr v0.9.1

```

```julia
using YAXArrays, NetCDF
using Downloads
using Glob
using Dates
using CFTime

# MWE. Sample files for creating the MWE. 
url = "https://esgf-world.s3.amazonaws.com/CMIP6/CMIP/CCCma/CanESM5/historical/r15i1p1f1/day/tasmin/gn/v20190306/tasmin_day_CanESM5_historical_r15i1p1f1_gn_18500101-20141231.nc"
filename = Downloads.download(url, "tasmin_day_CanESM5_historical_r15i1p1f1_gn_18500101-20141231.nc") # you pick your own path

# MWE hack. We copy the same file to mimic the loop over multiple files
cp(filename, "tasmin_day_CanESM5_historical_r15i1p1f1_gn_18500101-20141231_v2.nc")

# ==================================
# Beginning here, this is the actual MWE to optimize
files = glob("tasmin*.nc")

# To store filenames given different level of missing values
miss_frac = Vector{String}()
miss_over90 = Vector{String}()
miss_tot = Vector{String}()
miss_ok = Vector{String}()

for ifile in files

    ds = Cube(ifile)

    # MWE hack. This bit is just to optimize the code over a subset of the data
    # (the actual code will use the whole dataset)
    ds = ds[Time=DateTimeNoLeap(2000,1,1)..DateTimeNoLeap(2002,12,31)]

    # Total items
    tot_items = prod(size(ds))

    # Missing items
    # This is where it takes time. I was thinking that this would be parrallelized given that I provide the following ENV variable (but the process only use 1 core in the end) 
   # export JULIA_NUM_THREADS=8 (defined before I launch the script)
    missing_items = sum(map(x-> isnan(x), ds))

    # Fraction
    fraction = (missing_items / tot_items)

    if fraction > 0.5 && fraction < 0.9
        println("Found > 0.5 and < 0.9")
        push!(miss_frac, ifile)
    elseif fraction >= 0.9 && fraction < 1.0
        println("Found >= 0.9 and < 1.0")
        push!(miss_over90, ifile)
    elseif fraction == 1.00
        println("Found == 1.0!")
        push!(miss_tot, ifile)
    else
        push(miss_ok, ifile)
    end
end

```

@Fliks (perhaps you have an idea?)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 7, 2023, 4:59pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/2 "2023-12-07T16:59:19Z")

</div>

> [@Balinus](#):
>
> ```julia
> I was thinking that this would be parrallelized given that I provide the following ENV variable (but the process only use 1 core in the end) 
> # export JULIA_NUM_THREADS=8
> 
> ```

Giving Julia more threads does not automatically parallelize your code. You have to call parallel functions, or write your own multi-threaded code. For example, ThreadsX.jl has a parallelized version of `sum` that you can call, which will utilize multiple threads.

Note also that you will often get better speedup by parallelizing at a coarser level than a finer level — in particular, I would think about parallelizing your _outermost loop_, the `for ifile in files`, with [`@threads`](https://docs.julialang.org/en/v1/base/multi-threading/). If you do that, in order to avoid race conditions you will want to put a mutex lock around your `if fraction [...] end` condition at the end, so that only one thread calls `push!` at a time.

PS. Better to call `sum(isnan, ds)` or `count(isnan, ds)` rather than `sum(map(x-> isnan(x), ds))`, since the latter will first allocate an array of `isnan` values and _then_ sum it, whereas the former will allocate no intermediate arrays and make a single pass over the data.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 5:16pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/3 "2023-12-07T17:16:58Z")

</div>

Thanks!

I was under the impression that using @threads macro for the files would force me to create an array for `ds`, `tot_items`, `missing_items` and `fraction`, like:

```julia

# Code preallocating Arrays (ds, tot_items, etc) not shown

Threads.@threads for z in 1:length(files)

    ifile = files[z]
    
    ds[z] = Cube(ifile)
    
    tot_items[z] = prod(size(ds[z]))
    
    # Missing items
    missing_items[z] = sum(map(x-> isnan(x), ds[z]))
    
    # Fraction
    fraction[z] = (missing_items[z] / tot_items[z])
end

```

I will try your suggestion!

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [December 7, 2023, 5:24pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/4 "2023-12-07T17:24:22Z")

</div>

On a cluster think of using an Array Job . Your Julia code could count the NaNs in a given file. If you can run this code independently on each file then the array job can help.

ps. If it was possible to use grep to match NaNs I would suggest using grep…

[Slurm Workload Manager - Job Array Support (schedmd.com)](https://slurm.schedmd.com/job_array.html)

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 5:36pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/5 "2023-12-07T17:36:20Z")

</div>

Thanks!

Yes, I could indeed use a job array for the loop of the files. My hope was to distribute the file using a job array (or @threads loop) and then using multiple threads for the NaNs computations.

---

<div class="post-metadata">

**Author:** ![mike.ingold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike.ingold/32/203749_2.png) [@mike.ingold](https://discourse.julialang.org/u/mike.ingold)\
**Post date:** [December 7, 2023, 6:00pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/6 "2023-12-07T18:00:21Z")

</div>

What kinds of files are these? Plaintext or some binary format? If plaintext, depending on your exact goals, you might also consider just using a shell script with some combination of `wc`, `grep`, `sed`, `awk`, etc.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 6:11pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/7 "2023-12-07T18:11:25Z")

</div>

Indeed, I should have mentioned that those are netcdf files, which limits the access to the data to some librairies (here using YAXArrays, using NetCDF under-the-hood).

> **[Unidata | NetCDF](https://www.unidata.ucar.edu/software/netcdf/)**
>
> NetCDF is a set of software libraries and self-describing, machine-independent data formats that support the creation, access, and sharing of array-oriented scientific data.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [December 7, 2023, 6:48pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/8 "2023-12-07T18:48:04Z")

</div>

> [@Balinus](#):
>
> I was under the impression that using @threads macro for the files would force me to create an array for `ds`, `tot_items`, `missing_items` and `fraction`, like:

No, local variables within the loop body are already separate for each thread.

---

<div class="post-metadata">

**Author:** ![neophytedave](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/neophytedave/32/202273_2.png) [@neophytedave](https://discourse.julialang.org/u/neophytedave)\
**Post date:** [December 7, 2023, 8:11pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/9 "2023-12-07T20:11:31Z")

</div>

These are exactly my thoughts. Utilize various Linux utilities to do various part of this.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 8:57pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/10 "2023-12-07T20:57:18Z")

</div>

Yeah, I guess I could use `cdo` or `nco` and bash scripting, but I must admit I do not use shell scripting enough to do it quickly 🙂

---

<div class="post-metadata">

**Author:** ![mike.ingold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike.ingold/32/203749_2.png) [@mike.ingold](https://discourse.julialang.org/u/mike.ingold)\
**Post date:** [December 7, 2023, 8:59pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/11 "2023-12-07T20:59:16Z")

</div>

> [@Balinus](#):
>
> ```julia
> # Missing items
> missing_items[z] = sum(map(x-> isnan(x), ds[z]))
> 
> ```

I’m no allocations expert, but you might see a minor speed improvement here by replace this `sum(map(isnan,ds[z])` with `count(isnan, ds[z])`.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 7, 2023, 9:02pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/12 "2023-12-07T21:02:34Z")

</div>

I have not see major difference in allocations between `count(isnan,ds)`, `sum(isnan,ds)` and `sum(map(x-> isnan(x), ds))` after some testing. My guess is that the `map` function for an YAXArray is lazy by design? I could be wrong of course.

" Apply a function on every element of a datacube[¤](https://juliadatacubes.github.io/YAXArrays.jl/dev/examples/generated/UserGuide/applyfunctions/#apply-a-function-on-every-element-of-a-datacube)

The `map` function can be used to apply a function on every entry of a YAXArray without taking the dimensions into account. This will lazily register the mapped function which is applied when the YAXArray is either accessed or when more involved computations are made."

> **[Apply functions on YAXArrays - YAXArrays](https://juliadatacubes.github.io/YAXArrays.jl/dev/examples/generated/UserGuide/applyfunctions/#apply-a-function-on-every-element-of-a-datacube)**
>
> Yet Another XArray-like Julia Package

---

<div class="post-metadata">

**Author:** ![Fliks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fliks/32/2494_2.png) [@Fliks](https://discourse.julialang.org/u/Fliks)\
**Post date:** [December 8, 2023, 11:15am UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/13 "2023-12-08T11:15:35Z")

</div>

I think, that threading is not your main bottleneck. Your main bottleneck is that your code is running into a generic definition of sum which assumes fast random access to the data.  
When you replace your `missing_items` computation by

```julia
    @time missing_items_ts = mapCube(countnan, ds, indims=InDims("lon", "lat"), outdims=OutDims(), nthreads=1)

```

This computation takes the chunking of the underlying data into account and has to open every chunk only once.  
This is not going to be multithreaded, but it runs in half a second for the small dataset and in 14 seconds for the full examplary dataset for me while your code took 6 seconds for the small subset. You could then parallelize the computations via slurm along the number of files.  
Threading is not going to help much, because your problem is mainly IO limited.

For your full problem what is the chunking structure of the data?  
If a full lat lon chunk does not fit into memory you might have to do the computation in three steps along the longest chunking size. Or you could increase the max\_cache for the mapCube function to allow for more data to be loaded at the same time.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 8, 2023, 3:46pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/14 "2023-12-08T15:46:44Z")

</div>

Thanks!

Did you defined `countnan` somewhere?

---

<div class="post-metadata">

**Author:** ![Fliks](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fliks/32/2494_2.png) [@Fliks](https://discourse.julialang.org/u/Fliks)\
**Post date:** [December 8, 2023, 4:22pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/15 "2023-12-08T16:22:49Z")

</div>

Yes, sorry I forgot to include it here.

```julia
countnan(xout, xin) = xout .= count(isnan, xin)

```

This definition is needed so that YAXArrays is doing that computation in place and is going to reuse the temporary array for every time step.

---

<div class="post-metadata">

**Author:** ![Balinus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/balinus/32/243_2.png) [@Balinus](https://discourse.julialang.org/u/Balinus)\
**Post date:** [December 8, 2023, 4:39pm UTC](https://discourse.julialang.org/t/looping-over-a-thousand-files-and-checking-if-empty/107291/16 "2023-12-08T16:39:31Z")

</div>

> [@Fliks](#):
>
> `countnan(xout, xin) = xout .= count(isnan, xin)`

Awesome, thank you!

This is indeed 3x faster 🙂
