# How to accelerate the imfiter() operation?

**URL:** <https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965>\
**Category:** Signal and Image Processing\
**Created:** [August 18, 2023, 7:32pm UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965 "2023-08-18T19:32:37Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![RunguangLi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/runguangli/32/18012_2.png) [@RunguangLi](https://discourse.julialang.org/u/RunguangLi)\
**Post date:** [August 18, 2023, 7:32pm UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/1 "2023-08-18T19:32:37Z")

</div>

My task contains hundreds of thousands of convolution operations like the following one. Each cost 11s now. I wonder if anyone can help to accelerate it.

```julia
using ImageFiltering
@time imfilter(Float32, rand(Int8, 300, 300, 300), rand(Int8, 100, 100, 100));

```

> 11.034938 seconds (83 allocations: 3.212 GiB, 4.00% gc time)

One way that may be useful is to call GPU while the command

`imfilter(ArrayFireLibs(), rand(Int8, 300, 300, 300), rand(Int8, 100, 100, 100))` seems not to work as ArrayFire package cannot be installed by **add ArrayFire**.  
Any help is greatly appreciated.

---

<div class="post-metadata">

**Author:** ![RoyiAvital](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/royiavital/32/571_2.png) [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Post date:** [July 6, 2024, 9:57am UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/2 "2024-07-06T09:57:51Z")

</div>

Your case is using “images” with many channels.  
It might be that your best bet is using Deep Learning based convolution kernel.  
Though those are usually optimized to a small kernel, yet it is still worth trying.

---

<div class="post-metadata">

**Author:** ![brianguenter](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/brianguenter/32/29519_2.png) [@brianguenter](https://discourse.julialang.org/u/brianguenter)\
**Post date:** [July 6, 2024, 5:40pm UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/3 "2024-07-06T17:40:21Z")

</div>

The `imfilter` function is supposedly multithreaded but my 8 core machine runs at 24% utilitzation for your convolution example (on Julia v1.11).

Can your hundreds of thousands of convolutions be done in parallel? You might get better utilization this way. Unfortunately `imfilter` allocates a ton - this usually results in poor multithreading performance, since the garbage collector is not fully multithreaded.

There is an in place version of `imfilter!` but it allocates just marginally less than `imfilter`.

```julia
julia> @benchmark imfilter(Float32,$a,$b)
BenchmarkTools.Trial: 2 samples with 1 evaluation.
 Range (min … max): 3.246 s … 3.435 s ┊ GC (min … max): 4.08% … 9.22%
 Time (median): 3.340 s ┊ GC (median): 6.72%
 Time (mean ± σ): 3.340 s ± 133.388 ms ┊ GC (mean ± σ): 6.72% ± 3.64%

  █ █
  █▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ ▁
  3.25 s Histogram: frequency by time 3.43 s <

 Memory estimate: 3.19 GiB, allocs estimate: 117.

julia> @benchmark imfilter!($blank,$a,$b)
BenchmarkTools.Trial: 2 samples with 1 evaluation.
 Range (min … max): 3.023 s … 3.206 s ┊ GC (min … max): 4.82% … 9.61%
 Time (median): 3.115 s ┊ GC (median): 7.28%
 Time (mean ± σ): 3.115 s ± 129.296 ms ┊ GC (mean ± σ): 7.28% ± 3.39%

  █ █
  █▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁█ ▁
  3.02 s Histogram: frequency by time 3.21 s <

 Memory estimate: 3.09 GiB, allocs estimate: 114.

```

Given the size of your filter kernel `imfilter!` is most probably using the FFT in which case this [function](https://github.com/JuliaImages/ImageFiltering.jl/blob/66cf9d9a5d12208b0b855fe89a387a7957dff937/src/imfilter.jl#L843-L847) seems to be the source of the allocations:

```julia
function filtfft(A, krn)
    B = rfft(A)
    B .*= conj!(rfft(krn))
    irfft(B, length(axes(A, 1)))
end

```

If you could convince the authors of ImageFiltering.jl to make an in place version of this function, or write one yourself and submit a PR, you might get much better performance.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 6, 2024, 7:52pm UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/4 "2024-07-06T19:52:28Z")

</div>

> [@RoyiAvital](#):
>
> Your case is using “images” with many channels.  
> It might be that your best bet is using Deep Learning based convolution kernel.  
> Though those are usually optimized to a small kernel, yet it is still worth trying.

I’m not sure why you’re reviving an 11 months old question, but that `imfilter` is a 3D filter (i.e. three spatial dimensions) with a huge kernel, and in deep learning parlance with a single channel, so it seems unlikely that any deep learning library would be of much help.

---

<div class="post-metadata">

**Author:** ![RoyiAvital](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/royiavital/32/571_2.png) [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Post date:** [July 6, 2024, 10:57pm UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/5 "2024-07-06T22:57:18Z")

</div>

> [@GunnarFarneback](#):
>
> > [@RoyiAvital](#):
> >
> > Your case is using “images” with many channels.  
> > It might be that your best bet is using Deep Learning based convolution kernel.  
> > Though those are usually optimized to a small kernel, yet it is still worth trying.
> 
> but that `imfilter` is a 3D filter (i.e. three spatial dimensions) with a huge kernel, and in deep learning parlance with a single channel, so it seems unlikely that any deep learning library would be of much help.

I am not sure what you mean.  
If `imfilter` applies a 3D convolution in the case above, then the OP should use [`Conv3d`](https://pytorch.org/docs/stable/generated/torch.nn.Conv3d.html) (Borrowing the naming form PyTorch).

In my post I already mentioned that NN operators are usually optimized for small kernels. So in the case above they might not be useful.  
In light of that, I’m not sure what your comment is adding. Could you elaborate?

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 7, 2024, 7:44am UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/6 "2024-07-07T07:44:49Z")

</div>

Deep learning libraries are specialized for small kernels and multiple channels, neither of which is the case here. I mostly wanted to point out that the claim that this operation worked on many channels was a misunderstanding.

Addendum: Apologies for the inappropriate tone of my first message. I could have conveyed my points more professionally.

---

<div class="post-metadata">

**Author:** ![RoyiAvital](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/royiavital/32/571_2.png) [@RoyiAvital](https://discourse.julialang.org/u/RoyiAvital)\
**Post date:** [July 8, 2024, 7:28am UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/7 "2024-07-08T07:28:25Z")

</div>

Hi,

In the OP I see `rand(Int8, 300, 300, 300)` which is a 300 channels image.  
That what triggered me to suggest DL Libraries. I’d still give it a try.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [July 8, 2024, 7:50am UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/8 "2024-07-08T07:50:48Z")

</div>

It’s a matter of terminology but deep learning libraries distinguish between spatial dimensions and channels. A three-dimensional array could be two spatial dimensions and a channel dimension but in this case it isn’t since `imfilter` only works with spatial dimensions. Yes, you could append a singleton channel dimension and send it to a deep learning library with support for 3D convolutions and get a result, but it’s far from the use cases those libraries are optimized for.

---

<div class="post-metadata">

**Author:** ![roflmaostc](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/roflmaostc/32/30123_2.png) [@roflmaostc](https://discourse.julialang.org/u/roflmaostc)\
**Post date:** [July 8, 2024, 11:06am UTC](https://discourse.julialang.org/t/how-to-accelerate-the-imfiter-operation/102965/9 "2024-07-08T11:06:53Z")

</div>

Without testing, the fastest way for such a big kernel is FFT based.  
This also runs on CUDA, below a code snippet:

```julia
using NDTools, CUDA, FFTW

# untested 
conv(x, y; dims=(1,2)) = irfft(rfft(x, dims) .* rfft(y, dims), size(x, dims[1]), dims)

x = rand(Int8, 300, 300, 300)
y = rand(Int8, 100, 100, 100)) 

# if the kernel is large in space, you might need zero padding before 
# and cropping after to avoid circular wrap arounds with FFT based confs
round.(Int8, conv(x, select_region(y, x)))

```
