# Using BenchmarkTools with CUDAnative and CuArrays and running out of CPU or GPU memory

**URL:** <https://discourse.julialang.org/t/using-benchmarktools-with-cudanative-and-cuarrays-and-running-out-of-cpu-or-gpu-memory/27278>\
**Category:** General Usage\
**Tags:** cudanative, benchmarktools\
**Created:** [August 7, 2019, 5:41pm UTC](https://discourse.julialang.org/t/using-benchmarktools-with-cudanative-and-cuarrays-and-running-out-of-cpu-or-gpu-memory/27278 "2019-08-07T17:41:54Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![Josiah\_Slack](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josiah_slack/32/8319_2.png) [@Josiah\_Slack](https://discourse.julialang.org/u/Josiah_Slack)\
**Post date:** [August 7, 2019, 5:41pm UTC](https://discourse.julialang.org/t/using-benchmarktools-with-cudanative-and-cuarrays-and-running-out-of-cpu-or-gpu-memory/27278/1 "2019-08-07T17:41:54Z")

</div>

I’m doing some simple “get acquainted” experimenting with convolutions using DSP, CUDAnative, CUDAdrv and CuArrays. I’m creating random 3-d arrays - rand(Float32, N, N, N). I then create “device” versions of the arrays by calling cu(a)

> A = rand(Float32, N, N, N);  
> B = rand(Float32, N, N, N);  
> A\_d = cu(A);  
> B\_d = cu(B);

I’ve written a simple function to perform an convolution on a pair of arrays:

> function cuFFT(A, B)  
> C = conv(A, B)  
> finalize( C )  
> C =   
> end

Finally, I use BenchmarkTools’s @benchmark macro:

> @benchmark cuFFT($A\_d, $B\_d)

If I set N to, say, 64, Julia returns this error:

> ERROR: LoadError: CUFFTError(code 2, cuFFT failed to allocate GPU or CPU memory)

However if I set N 10 120, my script runs to completion.

When I originally posted my question, I was directly calling fft(). As I continued experimenting, I found that I was getting inconsistent results from run to run. However, I found that if I loaded DSP and called conv(), I was able to see consistent behavior, and the new puzzle that a larger N didn’t crash when a smaller N did. I also realized that I’d been assuming the problem was GPU memory, though the error message says “CPU or GPU”.

My question: is there some problem in the way that I’m calling DSP.conv(), or some setup that I need to do with BenchmarkTools?
