# Best way to call an OpenCL kernel with arguments of type CLArray

**URL:** <https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368>\
**Category:** GPU\
**Tags:** question, gpuarrays\
**Created:** [June 2, 2018, 1:55pm UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368 "2018-06-02T13:55:27Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [June 2, 2018, 1:55pm UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/1 "2018-06-02T13:55:28Z")

</div>

I would like to perform some computations on GPUs using `CLArray`s, since these support standard array syntax. However, I would like to call some existing OpenCL kernels on these `CLArray`s. Some possibilities that might seem to be natural (at least for me) do not work.

```julia
using OpenCL, CLArrays
device, ctx, queue = cl.create_compute_context()
mult_kernel = """
kernel void mult(global float const* a, global float* b)
{
  int gid = get_global_id(0);
  b[gid] = 2*a[gid];
}
"""
p = cl.Program(ctx, source=mult_kernel) |> cl.build!
mult_cl = cl.Kernel(p, "mult")

# using buffers: calling kernels works, but buffers do not support the array interface
a = rand(Float32, 50_000)
a_buff = cl.Buffer(Float32, ctx, (:r, :copy), hostbuf=a)
b_buff = cl.Buffer(Float32, ctx, :rw, length(a))
queue(mult_cl, size(a), nothing, a_buff, b_buff)
b = cl.read(queue, b_buff)
@show norm(b - 2a)

# calling queue with arguments of type CLArray throws an error
d_a = CLArray(a)
d_b = CLArray(similar(a))
queue(mult_cl, size(a), nothing, d_a, d_b)

# using gpu_call with a julia function works, but I would like to call an existing OpenCL kernel
function mult_julia(state, a, b)
  idx = GPUArrays.@linearidx a state
  @inbounds b[idx] = 2*a[idx]
end
gpu_call(mult_julia, d_a, (d_a, d_b))
mapreduce(x->x^2, +, d_b-2*d_a)

# calling gpu_call with an OpenCL kernel throws an error
gpu_call(mult_cl, d_a, (d_a, d_b))

```

What is the best way to call existing OpenCL kernels with `CLArray`s as arguments?

---

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [June 2, 2018, 5:16pm UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/2 "2018-06-02T17:16:37Z")

</div>

After browsing the source of CLArrays.jl and OpenCL.jl, I might have found a solution.

```julia
# this works
ctx = CLArrays.context(d_a)
queue = CLArrays.global_queue(d_a)
p = cl.Program(ctx, source=mult_kernel) |> cl.build!
mult_cl = cl.Kernel(p, "mult")
queue(mult_cl, size(d_a), nothing, pointer(d_a), pointer(d_b))
mapreduce(x->x^2, +, d_b-2*d_a)

```

Here, it is essential that the command queue `queue` and the context `ctx` are the corresponding ones of the `CLArray`s. Otherwise, I get `CLError(code=-38, CL_INVALID_MEM_OBJECT)`.

Nevertheless, I would like to know whether this approach works in general and whether there is some better possibility.

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [June 4, 2018, 9:41am UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/3 "2018-06-04T09:41:50Z")

</div>

I made a pr to have this integrated a bit nicer:  
[https://github.com/JuliaGPU/CLArrays.jl/pull/30](https://github.com/JuliaGPU/CLArrays.jl/pull/30)

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [June 4, 2018, 9:48am UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/4 "2018-06-04T09:48:52Z")

</div>

Your solution is also fine! 🙂

---

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [June 4, 2018, 9:56am UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/5 "2018-06-04T09:56:19Z")

</div>

Thank your very much, Simon!

In your PR, you wrote “_Note, that the caching of the functor is not very nice, so for repeated calls, one might want to do this part manually:_”  
So, for repeated calls, I should call `clfunc = CLFunction(f, _args, ctx)` only once and then use `clfunc(_args, global_size, threads)`, correct?

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [June 4, 2018, 9:58am UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/6 "2018-06-04T09:58:18Z")

</div>

Yes! Or benchmark the difference 😉 Would be interesting to know how bad the dictionary look up really is 😉  
Another side effect ist, that I’m not hashing the actual kernel string and instead just the function name + function argument types.

---

<div class="post-metadata">

**Author:** ![ranocha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ranocha/32/35588_2.png) [@ranocha](https://discourse.julialang.org/u/ranocha)\
**Post date:** [June 4, 2018, 9:59am UTC](https://discourse.julialang.org/t/best-way-to-call-an-opencl-kernel-with-arguments-of-type-clarray/11368/7 "2018-06-04T09:59:45Z")

</div>

Okay, thank you again. I will test it when I’m back at a machine running OpenCL…
