# Loading lots of images

**URL:** <https://discourse.julialang.org/t/loading-lots-of-images/17125>\
**Category:** Performance\
**Created:** [November 3, 2018, 8:55pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125 "2018-11-03T20:55:08Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 3, 2018, 8:55pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/1 "2018-11-03T20:55:08Z")

</div>

I have a lot of folders, each with a lot of images. All images have the same size.  
I want to compute the “average image” of each folder, that is, the pixel-wise average of the images in the folder.

With the `Images` package, the following code get the job done, where `files` is a vector with file names:

```
function average_image(files, sz)
    image_sum = zeros(Gray{Float64}, sz)

    for file in files
        image_sum += load(file)
    end

    return(image_sum / length(files))
end

```

Unfortunately, this is slow and use a lot of memory.  
I suspect that every image is loaded into its own chunk of memory and soon after discarded by the garbage collector. But since they all have the same size I wonder if it’s possible to load each image into the same chunk of memory?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [November 3, 2018, 10:56pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/2 "2018-11-03T22:56:30Z")

</div>

Take a look at the `@code_warntype` of your function - there’s a bunch of type conversion happening and thus a bunch of `Any`. In particular, `image_sum += load(..)` and `image_sum / length(files)` are the offenders - broadcasting those functions should make it faster (i.e., `.+=` and `./` respectively), but I don’t think you’re going to be able to remove all of them because of that `load`.

It’s difficult to give concrete advice though - do you have some benchmarking results you can compare to, maybe some small subset of the images as a testing ground?

---

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 4, 2018, 1:42pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/3 "2018-11-04T13:42:07Z")

</div>

Thanks! I feel pretty stupid for not `@code_warntype`ing.

One thing I find peculiar is that if I follow your advice with `.+=` and `./` I reduce the memory usage, but `@code_warntype` still gives an `Any` type in an intermediate result, although the final result has a known type.  
I can remove this explicitly type annotating the `load`.

But I suppose it’s difficult for `load` to know the type of the file it loads?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [November 4, 2018, 1:54pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/4 "2018-11-04T13:54:03Z")

</div>

> [@robertdj](#):
>
> But I suppose it’s difficult for `load` to know the type of the file it loads?

Precisely! At compile time the return type of load cannot really be known for the general case - maybe there are some hints to give in that function curtesy of Images.jl? You’ll have to check the docs on that.

Those changes alone should already be pretty much the best you can do without too much work, I think.

---

<div class="post-metadata">

**Author:** ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)\
**Post date:** [November 4, 2018, 1:57pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/5 "2018-11-04T13:57:56Z")

</div>

Also, if you ever need more than just the mean (via sum), then check [this out](https://github.com/joshday/OnlineStats.jl/issues/134). I needed something similar too…

---

<div class="post-metadata">

**Author:** ![robertdj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/robertdj/32/103_2.png) [@robertdj](https://discourse.julialang.org/u/robertdj)\
**Post date:** [November 5, 2018, 7:59pm UTC](https://discourse.julialang.org/t/loading-lots-of-images/17125/6 "2018-11-05T19:59:57Z")

</div>

I’m having one more problem: I would like to run the computation for the folders in “parallel” and reading the [documentation about channels](https://docs.julialang.org/en/v1/manual/parallel-computing/) I see that channels are well suited for such an I/O intensive task.

However, since I have a lot folders, I would like to control the number of “concurrent” tasks. The documentation I link to above have an example where sleeping processes run 4 at a time, but I cannot figure out how to adapt that to my situation.

With an `average_image` function like in my first post I have a wrapper that saves the output:

```
save_average(dir, sz)
    files = joinpath.(dir, readdir(dir))
    avgimg = average_image(files, sz)

    savename = string(dir, ".png")
    save(savename, avgimg)
end

```

Processing all folders at once as explained in [this question](https://discourse.julialang.org/t/example-of-async-distributed-loading-into-a-single-shared-channel/11815), which I fear is too agressive:

```
@sync for dir in dirs
    @async save_average(dir, sz)
end

```
