# Optimizing the use of Blocks, Threads vs. Array Indexing

**URL:** <https://discourse.julialang.org/t/optimizing-the-use-of-blocks-threads-vs-array-indexing/7537>\
**Category:** GPU\
**Created:** [December 5, 2017, 3:59pm UTC](https://discourse.julialang.org/t/optimizing-the-use-of-blocks-threads-vs-array-indexing/7537 "2017-12-05T15:59:57Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [December 6, 2017, 10:12am UTC](https://discourse.julialang.org/t/optimizing-the-use-of-blocks-threads-vs-array-indexing/7537/2 "2017-12-06T10:12:54Z")

</div>

> Is the above code optimal

Well, you can just benchmark it, right?

Given this slightly modified example which allows me to quickly shift indexes around (with the 1024 thread limit of my GPU):

```julia
function kernel(x)
    c = blockIdx().x
    b = blockIdx().y
    a = threadIdx().x
    x[a, b, c] = x[a, b, c] + 1
    return
end

dx = CuArray{Float32,3}(32, 32, 32)
@cuda ((32, 32), 32) kernel(dx)

```

Now let’s profile this:

```julia
$ nvprof julia wip.jl
==15727== NVPROF is profiling process 15727, command: julia wip.jl
==15727== Profiling application: julia wip.jl
==15727== Profiling result:
            Type Time(%) Time Calls Avg Min Max Name
 GPU activities: 100.00% 7.3280us 1 7.3280us 7.3280us 7.3280us ptxcall_kernel_63191

```

Now swap the indexing around, and benchmark again. You’ll see that the situation where the innermost index is the fastest evolving one, triggers the global memory coalescing as I described in detail in [CuArray is Row Major or Column Major? - #2 by maleadt](https://discourse.julialang.org/t/cuarray-is-row-major-or-column-major/7402/2)

> How much of a performance gain/hit will my decision affect.

On my GPU (first generation Titan), there’s a 2x penalty by indexing inefficiently. Of course, in the presence of other operations, and with a higher occupancy, this penalty might be significantly lower.

On an unrelated note, please mark threads as "solved’ if they’ve answered your questions.  
Makes it easier to maintain overview of the GPU category 🙂

---

_[View the full topic](https://discourse.julialang.org/t/optimizing-the-use-of-blocks-threads-vs-array-indexing/7537)._
