# Computational performance when broadcasting a convolution layer

**URL:** <https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587>\
**Category:** Machine Learning\
**Tags:** performance, flux\
**Created:** [August 18, 2021, 7:25am UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587 "2021-08-18T07:25:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![MattSainsbury-Dale](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattsainsbury-dale/32/221714_2.png) [@MattSainsbury-Dale](https://discourse.julialang.org/u/MattSainsbury-Dale)\
**Post date:** [August 18, 2021, 7:25am UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/1 "2021-08-18T07:25:32Z")

</div>

For 2D images, the usual input format for `Conv` layers is `Array{Float32, 4}`. I’m currently experimenting with broadcasting a `Conv` layer over a `Vector{Array{Float32, 4}}`, however I have found the performance to be significantly worsened by this approach.

```julia
using Flux
using BenchmarkTools

const w = 16 # image width
const h = 16 # image height
CNN = Conv((5, 5), 1 => 4) # arbitrarily chosen architecture

X = rand(Float32, w, h, 1, 1000)
@btime CNN($X)
# 1.606 ms (58 allocations: 4.45 MiB)

X = [rand(Float32, w, h, 1, 1) for _ ∈ 1:1000]
@btime CNN.($X)
# 14.195 ms (56000 allocations: 62.69 MiB)

```

This performance discrepancy increases as the architecture complexity increases. Does anyone know why there is a performance discrepancy? (There are reasons why I would want to use this structure.)

---

<div class="post-metadata">

**Author:** ![contradict](https://avatars.discourse-cdn.com/v4/letter/c/ac91a4/32.png) [@contradict](https://discourse.julialang.org/u/contradict)\
**Post date:** [August 18, 2021, 3:39pm UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/2 "2021-08-18T15:39:56Z")

</div>

The difference in allocations probably accounts for most of the discrepancy. In the broadcast case, all the intermediates needed must be re-allocated for each application of the model. In the other case, you get a few big allocations. In addition, it is often the case on modern systems (especially when you use a GPU, but even on CPUs) that one big matrix multiply is faster than many small matrix multiplies.

One possible solution is to use [RecursiveArrayTools.jl](https://github.com/SciML/RecursiveArrayTools.jl) which would allow you to view the data as a vector of arrays or as one big array as needed.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [August 18, 2021, 5:07pm UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/3 "2021-08-18T17:07:49Z")

</div>

> [@MattSainsbury-Dale](#):
>
> (There are reasons why I would want to use this structure.)

I assume in practice you would be using minibatches instead of batches of size 1, so the effect should be less pronounced. It might not hurt to articulate why this is desired though, as other than memory liminations I can’t think of any reason to use a smaller batch size.

---

<div class="post-metadata">

**Author:** ![MattSainsbury-Dale](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattsainsbury-dale/32/221714_2.png) [@MattSainsbury-Dale](https://discourse.julialang.org/u/MattSainsbury-Dale)\
**Post date:** [August 19, 2021, 2:59am UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/4 "2021-08-19T02:59:12Z")

</div>

Thanks very much for the comments.

As TouchSir notes, some context may be helpful (I wanted to keep the MWE as simple as possible, but that simplicity has probably been a hindrence).

At m locations, we have n\_i, i = 1, \dots, m, univariate (single channel) 16 \times 16 input images. In the original example, we have a single image at each location, so n\_i = 1: However, in practice, n\_i can be any positive integer and is not necessarily constant. For each location, I need to apply a CNN to each of the n\_i images, and then aggregate the resulting n\_i outputs associated with a given location.

Some code which should add clarity. If we have m = 1000 locations and between 20 and 40 images at each location, then I would do the following (including an attempt at using `RecursiveArrayTools.jl`, as per contradict’s suggestion):

```julia
using RecursiveArrayTools

m = 1000
n = rand(20:40, m)
X = [rand(Float32, w, h, 1, n[i]) for i ∈ eachindex(n)]

@btime CNN.($X)
# 177.813 ms (58225 allocations: 190.71 MiB)

A = ArrayPartition(X...)
@btime CNN.($A.x)
# 169.872 ms (58003 allocations: 190.14 MiB)

X = VectorOfArray(X)
@btime CNN.($X.u) 
# 172.439 ms (58103 allocations: 190.13 MiB)

```

(I may not have implemented `RecursiveArrayTools.jl` as contradict had intended; please let me know if this is the case.) For comparison with a similar amount of data using the traditional approach of feeding in a single large array:

```julia
X = rand(Float32, w, h, 1, 30 * 1000)
@btime CNN($X)
# 73.520 ms (58 allocations: 131.89 MiB)

```

I think this supports contradict’s point, that the poor performance comes from using many small arrays rather than one large array. Perhaps there is a better approach than broadcasting over a vector of arrays? I have looked into [`DepthwiseConv`](https://fluxml.ai/Flux.jl/stable/models/layers/#Flux.DepthwiseConv) layers, but I don’t think it’s exactly what I need, and it doesn’t seem to work on the GPU.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [August 19, 2021, 3:53am UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/5 "2021-08-19T03:53:12Z")

</div>

Thanks for the explanation. AIUI, you just need to be able to aggregate the CNN outputs for each location? If so, then it should be sufficient to run the model over the entire batch and _then_ do the aggregation. This assumes there are no layers like batchnorm present which break the independence of samples in the batch. It’s not the end of the world if you do have those layers either, but there will likely be more work involved (less so if you’re only concerned about inference).

---

<div class="post-metadata">

**Author:** ![MattSainsbury-Dale](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mattsainsbury-dale/32/221714_2.png) [@MattSainsbury-Dale](https://discourse.julialang.org/u/MattSainsbury-Dale)\
**Post date:** [August 24, 2021, 1:11am UTC](https://discourse.julialang.org/t/computational-performance-when-broadcasting-a-convolution-layer/66587/6 "2021-08-24T01:11:45Z")

</div>

Hi, sorry for the delayed response.

> [@ToucheSir](#):
>
> AIUI, you just need to be able to aggregate the CNN outputs for each location? If so, then it should be sufficient to run the model over the entire batch and _then_ do the aggregation.

Yes, that’s right. Appreciate the suggestion, and I spent some time exploring it. I did some tests on the run-time and memory performance using the “Array” method, which applies the CNN to one large array of images, and the “VecArray” method, which broadcasts the CNN over a vector of arrays of images. I was interested to see how performance changes with

- the size of the CNN, in terms of the number of trainable parameters,
- m, the number of “locations”; that is, the number of elements in the vector that we are broadcasting over, and
- n\_i, the number of images at each location; that is, the number of images stored in each element of the vector.

I looked at two CNNs, one with ~100 parameters, and another with ~420,000 parameters. The large CNN is reflective of the size of network I am using in my application. I used all combinations of m \in \{16, 32, 64, 128\} and n \_i \in \{1, 10, 25, 50, 75, 100\}. (These values of m are reflective of typical mini-batch sizes.)

**Run-times:** Using the small CNN yields a large discrepancy between run-times, while there is little difference when using the large CNN:

_Small CNN (~100 parameters):_

![CNN1_time](https://global.discourse-cdn.com/julialang/original/3X/6/7/6794bffe9194b491fe207bd686166bf48650e0d5.png)

_Large CNN (~420,000 parameters):_

![CNN2_time](https://global.discourse-cdn.com/julialang/original/3X/a/9/a995d4fc26e079ffc7ccb2e37231d48edc2b58f2.png)

**Memory usage:** there appears to be a constant penalty to using the Vector{Array} approach, which is fixed for a given value of m (i.e., it does not change with n\_i):

_Small CNN (~100 parameters):_

![CNN2_memory](https://global.discourse-cdn.com/julialang/original/3X/5/7/57f6847031a88be6e6e7b0a9ebdc25ed5ee9e7f9.png)

_Large CNN (~420,000 parameters):_

![CNN1_memory](https://global.discourse-cdn.com/julialang/original/3X/7/2/72022c230efaafd8adc24172fa09567904337175.png)

Note that this was done on the CPU: The results may change on the GPU. The results suggest that, for the large CNN, there isn’t a massive difference in terms of run time and memory usage when n\_i is “large enough” (e.g., for n\_i = 50).

I should note that I made a mistake by looking at the case of m = 1000 and n\_i = 1 in my original question. In this configuration, applying a CNN to an array and then aggregating is indeed much better than broadcasting over a vector of arrays, but this is not practically relevant for my application.

Thanks again for your helpful comments!
