# Is sharedmemory really accelerates GPU kernel?

**URL:** <https://discourse.julialang.org/t/is-sharedmemory-really-accelerates-gpu-kernel/123343>\
**Category:** Specific Domains\
**Tags:** gpu\
**Created:** [December 2, 2024, 7:17am UTC](https://discourse.julialang.org/t/is-sharedmemory-really-accelerates-gpu-kernel/123343 "2024-12-02T07:17:06Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [December 2, 2024, 1:39pm UTC](https://discourse.julialang.org/t/is-sharedmemory-really-accelerates-gpu-kernel/123343/2 "2024-12-02T13:39:13Z")

</div>

Shared memory is not going to always improve performance. For one, it may lower occupancy as it’s a shared resource limiting how many threads can be launched. But also, you seem to be using it here to simply cache accesses to read-only arrays. Modern GPUs are much better at automatically caching such reads, which may explain why shared memory doesn’t help here. It is still very relevant as a communication mechanism between threads, e.g., to implement a reduction.

If you want to be sure, run these two kernels under NSight Compute, which can show you accurately how memory is accessed and cached:

 ![image](https://global.discourse-cdn.com/julialang/original/3X/4/4/445ee252dd32d5b2507186b1db396a4505bfc323.png)

---

_[View the full topic](https://discourse.julialang.org/t/is-sharedmemory-really-accelerates-gpu-kernel/123343)._
