# Are SharedArrays bypassing cache?

**URL:** <https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269>\
**Category:** Julia at Scale\
**Created:** [February 6, 2020, 8:15pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269 "2020-02-06T20:15:25Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 6, 2020, 8:15pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269/1 "2020-02-06T20:15:25Z")

</div>

Hi, I am testing parallel computations based on multithreading/multiprocessing. Aimed at iterative solvers for PDEs I am interested in large loops. An interesting benchmark in this respect is the “Schönauer vector triad”, see e.g. the [benchmarking site of Georg Hager](https://blogs.fau.de/hager/archives/tag/benchmarking).

![](https://global.discourse-cdn.com/julialang/original/3X/a/f/afb055312a9f892804a5cc8b753b223f4e4a1091.png)

([Here](https://github.com/j-fu/julia-tests/tree/master/parallel) is the generating code)

From performing this test in the scalar case I see a striking performance difference between shared and normal arrays. For large arrays, the GFlop/s rates converge, in this case, for normal arrays the performance is limited by memory access due to exhaustion of the L3 cache. This lets me conclude that SharedArrays bypass the cache completely, which IMHO would be understandable as there needs to be some way to keep data coherent. I also suspect that this essentially due to the design of POSIX shared memory on the OS level and that Julia cannot do much about it. I googled for more evidence on this, but didn’t find any reasonable source.

Am I missing something ?

Please see also the MWE:

```julia
using SharedArrays
using BenchmarkTools

function vtriad(N,a,b,c,d)
    @inbounds @fastmath for i=1:N
        d[i]=a[i]+b[i]*c[i]
    end
end

function runtest(N)
    a = rand(N)
    b = rand(N)
    c = rand(N)
    d = rand(N)
    
    @btime vtriad($N,$a,$b,$c,$d)
    
    sa = SharedArray(a)
    sb = SharedArray(b)
    sc = SharedArray(c)
    sd = SharedArray(d)

    @btime vtriad($N,$sa,$sb,$sc,$sd)
end

runtest(1000)

```

---

<div class="post-metadata">

**Author:** ![j-fu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/j-fu/32/11373_2.png) [@j-fu](https://discourse.julialang.org/u/j-fu)\
**Post date:** [February 7, 2020, 1:23pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269/2 "2020-02-07T13:23:47Z")

</div>

Update: surpisingly, `@avx` makes a difference here, thanks to @tbeason for the hint:

[![](https://global.discourse-cdn.com/julialang/original/3X/2/d/2d046585d2ae1958256b841da22ccff58fcddd8d.png) ](https://raw.githubusercontent.com/j-fu/julia-tests/master/parallel/results/shared-vs-normal-arrays-v2.png)

This would mean that optimization for shared arrays on the Julia side matters, and the OS is not at play here…

Still need to check if the computed results are correct, though (after some admin homework ☹ )

Here is the updated MWE:

```julia
using SharedArrays
using BenchmarkTools
using LoopVectorization

function vtriad(N,a,b,c,d)
    @avx for i=1:N
        d[i]=a[i]+b[i]*c[i]
    end
end

function runtest(N)
    a = rand(N)
    b = rand(N)
    c = rand(N)
    d = rand(N)
    
    @btime vtriad($N,$a,$b,$c,$d)
    
    sa = SharedArray(a)
    sb = SharedArray(b)
    sc = SharedArray(c)
    sd = SharedArray(d)

    @btime vtriad($N,$sa,$sb,$sc,$sd)
end

runtest(1000)

```

---

<div class="post-metadata">

**Author:** ![Ralph\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ralph_smith/32/10344_2.png) [@Ralph\_Smith](https://discourse.julialang.org/u/Ralph_Smith)\
**Post date:** [February 7, 2020, 1:55pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269/3 "2020-02-07T13:55:53Z")

</div>

This is very surprising. Perhaps @Elrod could explain how `@avx` is working here, and say whether there is a possible data race.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [February 7, 2020, 6:50pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269/4 "2020-02-07T18:50:17Z")

</div>

I think the problem is [here](https://github.com/JuliaLang/julia/blob/2d5741174ce3e6a394010d2e470e4269ca54607f/stdlib/SharedArrays/src/SharedArrays.jl#L507).

Shouldn’t the `getindex` and `setindex!` definitions be `@propogate_inbounds`?

The `@code_native` shows it isn’t vectorized, and that there are bounds checks in the hot loop.

`@avx` doesn’t do bounds checks, which is why it was as fast as normal.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [February 7, 2020, 6:51pm UTC](https://discourse.julialang.org/t/are-sharedarrays-bypassing-cache/34269/5 "2020-02-07T18:51:58Z")

</div>

> [@Elrod](#):
>
> Shouldn’t the `getindex` and `setindex!` definitions be `@propogate_inbounds` ?

👍
