# Thread local storage is not implemented

**URL:** <https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656>\
**Category:** GPU\
**Created:** [September 15, 2020, 7:11pm UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656 "2020-09-15T19:11:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![barrettp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barrettp/32/18259_2.png) [@barrettp](https://discourse.julialang.org/u/barrettp)\
**Post date:** [September 15, 2020, 7:11pm UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656/1 "2020-09-15T19:11:09Z")

</div>

It would be nice to have “thread local storage” implemented at some point. It looks like it might be necessary for recursive kernel calls. I’m not really sure as I’m still rather new to CUDA programming.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [September 17, 2020, 9:21am UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656/2 "2020-09-17T09:21:49Z")

</div>

PTX’s local memory is backed by global memory, and is thus slow. Why not use StaticArrays for fast thread-local memory? Is there a particular reason you need the former?

---

<div class="post-metadata">

**Author:** ![barrettp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barrettp/32/18259_2.png) [@barrettp](https://discourse.julialang.org/u/barrettp)\
**Post date:** [September 19, 2020, 4:50pm UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656/3 "2020-09-19T16:50:38Z")

</div>

I am currently using static arrays. I’m still new to CUDA programming, so I’ll have to investigate this some more. Thanks.

---

<div class="post-metadata">

**Author:** ![Shuhua](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shuhua/32/27618_2.png) [@Shuhua](https://discourse.julialang.org/u/Shuhua)\
**Post date:** [February 21, 2021, 7:37am UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656/4 "2021-02-21T07:37:44Z")

</div>

> [@maleadt](#):
>
> Why not use StaticArrays for fast thread-local memory?

Hi, @maleadt, is it guaranteed that a static array is backed by thread-local registers in CUDA.jl? In NVIDIA’s [documentation](https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/index.html#local-memory), there is

> Local memory is so named because its scope is local to the thread, not because of its physical location. In fact, local memory is off-chip. Hence, access to local memory is as expensive as access to global memory.

And also

> Automatic variables that are likely to be placed in local memory are large structures or arrays that would consume too much register space and arrays that the compiler determines may be indexed dynamically.

What does “the compiler determines may be indexed dynamically” refer to in Julia? Let’s consider the following example.

```julia
using CUDA, StaticArrays

function kernel()
    sa = SA_F32[1, 2, 3, 4, 5]
    s = 0.0f32
    for i in eachindex(sa)
        s += sa[i]
    end
    @cuprintf("s = %f", s)
    nothing
end

@cuda kernel()

```

Is the static array `sa` dynamically indexed above (since `i` is not a constant)?

Besides, how can we inspect the generated code by CUDA.jl to confirm whether we are using registers or _local memory_ for the thread-local arrays?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [February 22, 2021, 12:53pm UTC](https://discourse.julialang.org/t/thread-local-storage-is-not-implemented/46656/5 "2021-02-22T12:53:50Z")

</div>

StaticArrays are implemented as structs containing tuples, so it’ll be using registers and not PTX’s local memory. You can inspect generated code using `@device_code_ptx`.
