# FFT'ing lots of short vectors on GPU

**URL:** <https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529>\
**Category:** GPU\
**Tags:** fftw\
**Created:** [January 31, 2022, 6:58pm UTC](https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529 "2022-01-31T18:58:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![wsshin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wsshin/32/360_2.png) [@wsshin](https://discourse.julialang.org/u/wsshin)\
**Post date:** [January 31, 2022, 6:58pm UTC](https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529/1 "2022-01-31T18:58:57Z")

</div>

I am wondering if there is a way to Fourier-transform lots of short vectors on GPU.

I have more than 10,000 vectors. The lengths of all vectors are the same and around `2^12`. Each vector is not long enough to get a significant performance boost by performing FFT on GPU. For example, running the following code

```julia
using CUDA, FFTW

n = 12
N = 2^n

v = rand(Complex{Float64}), N)
w = copy(v)
F = plan_fft(v)
@btime $w .= $F * $v

vg = cu(v)
wg = copy(vg)
Fg = plan_fft(vg)
@btime $wg .= $Fg * $vg

```

produces

```julia-repl
  24.818 µs (2 allocations: 64.05 KiB)
  9.964 µs (20 allocations: 960 bytes)

```

so GPU is only 2–3 times faster. In comparison, for a vector of length `2^20`, which is much longer than `2^12` tested above, I get

```julia-repl
  39.214 ms (2 allocations: 16.00 MiB)
  15.603 µs (20 allocations: 960 bytes)

```

so GPU is 3-orders-of-magnitude faster than CPU.

The above examples demonstrate that in order to get a significant performance boost on GPU, the amount of data needs to be sufficiently large. In my case, each vector is short but there are lots of them, so the total amount of data to FFT seems sufficiently large. I am wondering if there is a clever way to take advantage of my situation to perform FFT fast on GPU.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [January 31, 2022, 8:37pm UTC](https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529/3 "2022-01-31T20:37:04Z")

</div>

A side comment / question: Don’t you want to use [`CUDA.CUFFT`](https://cuda.juliagpu.org/stable/usage/array/#FFTs) (instead of FFTW) for FFTs on GPU? Or is it perhaps automatically used for `CuArrays`?

---

<div class="post-metadata">

**Author:** ![wsshin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wsshin/32/360_2.png) [@wsshin](https://discourse.julialang.org/u/wsshin)\
**Post date:** [January 31, 2022, 9:17pm UTC](https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529/4 "2022-01-31T21:17:31Z")

</div>

Actually I’m not sure exactly how it works, but if you import both `CUDA` and `FFTW`, the `AbstractFFT` functions implemented in `CUDA.CUFFT` are exposed somehow:

```julia-repl
julia> using CUDA, FFTW

julia> methods(plan_fft)
# 8 methods for generic function "plan_fft":
[1] plan_fft(X::CuArray{T, N}, region) where {T<:Union{ComplexF32, ComplexF64}, N} in CUDA.CUFFT at /home/gridsan/WO28768/.julia/packages/CUDA/sCev8/lib/cufft/fft.jl:296
...

```

---

<div class="post-metadata">

**Author:** ![wsshin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/wsshin/32/360_2.png) [@wsshin](https://discourse.julialang.org/u/wsshin)\
**Post date:** [January 31, 2022, 9:28pm UTC](https://discourse.julialang.org/t/ffting-lots-of-short-vectors-on-gpu/75529/5 "2022-01-31T21:28:22Z")

</div>

One potential method I hoped possible was to create a 2¹²-by-10,000 matrix `V` whose columns are the vectors I want to FFT. Because `F = plan_fft(...)` creates an operator that can be represented by a 2¹²-by-2¹² matrix, I hoped `F * V` to work, but it didn’t.

I suspect in principle `F * V` could be implemented such that it can be performed fast on GPU for a GPU matrix `V`, but I’m not sure. If this is possible and seems useful to other people, I may want to create a GitHub issue requesting this feature.
