# Fastest possible trilinear interpolation with \`Interpolations.jl\`

**URL:** <https://discourse.julialang.org/t/fastest-possible-trilinear-interpolation-with-interpolations-jl/71184>\
**Category:** Performance\
**Tags:** multithreading, interpolations, threads\
**Created:** [November 9, 2021, 7:07am UTC](https://discourse.julialang.org/t/fastest-possible-trilinear-interpolation-with-interpolations-jl/71184 "2021-11-09T07:07:25Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Chiil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chiil/32/27474_2.png) [@Chiil](https://discourse.julialang.org/u/Chiil)\
**Post date:** [November 9, 2021, 7:07am UTC](https://discourse.julialang.org/t/fastest-possible-trilinear-interpolation-with-interpolations-jl/71184/1 "2021-11-09T07:07:25Z")

</div>

I am working on developing a machine learning application in which two instances of our CFD code work simultaneously at different resolutions and data is passed between the two instances. For this, I need to have fastest possible interpolations. At the moment I work with `Interpolations.jl` but I find the performance insufficient for my application and the interpolations are a huge bottleneck. Due to the absense of thread support in `Interpolations.jl`, I am launching the interpolations of different fields on separate threads. I am not sure whether I am using `Interpolations.jl` to its maximum extent as the benchmarks on the site are very fast. How can I speed up the code below that is a minimal example of my use case? I have a factor 2 or 3 ratio in grid points per dimension so I could potentially make use of repeating patterns.

```julia
using Interpolations
using BenchmarkTools

function interpolate_test!(a_hi, a_lo, x_hi, x_lo)
    interp_a = interpolate((x_lo, x_lo, x_lo), a_lo, (Gridded(Linear()), Gridded(Linear()), Gridded(Linear())))
    a_hi[:, :, :] .= interp_a(x_hi, x_hi, x_hi)
end

n_hi = 256
n_lo = 128

dx_hi = 1//n_hi
dx_lo = 1//n_lo

x_hi = 1//2*dx_hi:dx_hi:1
x_lo = -1//2*dx_lo:dx_lo:1+1//2*dx_lo

a0_lo = rand(n_lo+2, n_lo+2, n_lo+2)
a0_hi = zeros(n_hi, n_hi, n_hi)

a1_lo = rand(n_lo+2, n_lo+2, n_lo+2)
a1_hi = zeros(n_hi, n_hi, n_hi)

a2_lo = rand(n_lo+2, n_lo+2, n_lo+2)
a2_hi = zeros(n_hi, n_hi, n_hi)

a3_lo = rand(n_lo+2, n_lo+2, n_lo+2)
a3_hi = zeros(n_hi, n_hi, n_hi)

@btime begin
    @sync begin
        Threads.@spawn interpolate_test!(a0_hi, a0_lo, x_hi, x_lo)
        Threads.@spawn interpolate_test!(a1_hi, a1_lo, x_hi, x_lo)
        Threads.@spawn interpolate_test!(a2_hi, a2_lo, x_hi, x_lo)
        Threads.@spawn interpolate_test!(a3_hi, a3_lo, x_hi, x_lo)
    end
end

```

---

<div class="post-metadata">

**Author:** ![maxfreu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maxfreu/32/17468_2.png) [@maxfreu](https://discourse.julialang.org/u/maxfreu)\
**Post date:** [November 9, 2021, 7:47am UTC](https://discourse.julialang.org/t/fastest-possible-trilinear-interpolation-with-interpolations-jl/71184/2 "2021-11-09T07:47:40Z")

</div>

In case you just need up- or downsampling you can use the code in NNlib or NNlibCUDA (upsample\_trilinear). It works on GPU and the CPU code is threaded.
