# ANN: Strided.jl , a high-level dense array acceleration package

**URL:** <https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434>\
**Category:** General Usage\
**Created:** [October 17, 2018, 8:45am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434 "2018-10-17T08:45:42Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [October 17, 2018, 8:45am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/1 "2018-10-17T08:45:42Z")

</div>

Dear Julia users,

I have recently registered a package called [Strided.jl](https://github.com/Jutho/Strided.jl) that can speed up certain tasks with dense arrays. Its primary use is when you combine arrays whose memory layout is strided but with strides that are not monotoneously increasing, e.g. the transpose/adjoint of an array, in e.g. a map or broadcast operation.

The main functionality of the package is obtained by using the macro `@strided` which acts as a binding contract that any arrays in the ensuing expression will have a strided memory layout. It then yields lazy views, lazy `permutedims` and overloads most of broadcasting, map(reduce) etc to give a speedup by using an efficient blocking strategy to loop through the different arrays.

Since these implementations are also multithreaded (using `Threads.@threads`) there might even be a speedup if all your arrays are just plain `Array`s with monotoneously increasing strides, provided you have `Threads.nthreads() > 1`.

See the examples and README for further information. Beware that bugs can certainly be possible!

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [October 17, 2018, 11:50am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/2 "2018-10-17T11:50:21Z")

</div>

Looks great!

1 Is this meant to prototype ideas for future inclusion in Base julia (without macros) at some point, or is inclusion fundamentally impossible?

2 Is this interesting only for large arrays, or is there also a benefit in small to mid-sized arrays?

---

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [October 17, 2018, 12:58pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/3 "2018-10-17T12:58:12Z")

</div>

1. That is not up to me, it’s not how this was intended but if the core developers consider this useful, then I certainly want to contribute.

2. There is some overhead, which might be significant for small arrays. I would say, it’s not hard to benchmark a particular use case with and without `@strided` in front.

---

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [October 17, 2018, 1:08pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/4 "2018-10-17T13:08:19Z")

</div>

Warning: there certainly seem to be a few bugs, i.e. with the result of e.g. a simple reduction not being thrustworthy. I’ll have to look into this a bit further. I have mostly been using it for `map!` and related operations (`axpy!`, `permutedims!`, …), but not for reductions.

---

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [October 17, 2018, 9:51pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/5 "2018-10-17T21:51:52Z")

</div>

Strided v0.1.2 has just been tagged (and merged into metadata), which fixes the bug with certain (map)reductions. Notice though that currently, Strided.jl does not offer any benefits (in particular, not multithreading) for complete reductions, i.e. `(map)reduce` without a `dims` keyword, or with the default value `dims=:`. I will consider adding this soon, but it requires a somewhat different implementation.

---

<div class="post-metadata">

**Author:** ![marius311](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/marius311/32/3953_2.png) [@marius311](https://discourse.julialang.org/u/marius311)\
**Post date:** [May 21, 2019, 10:50pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/6 "2019-05-21T22:50:42Z")

</div>

This is really great, thanks! I notice I can get speedups of several even for things like simple pointwise multiplication (which I might have though was memory-limited), e.g. with 8 threads on a NERSC Cori [Haswell](https://www.nersc.gov/users/computational-systems/cori/configuration/cori-phase-i/) node:

```julia
using Strided
using BenchmarkTools
A = randn(256,256);
B = similar(A);
@btime $B .= ($A .* $A);
# 40.658 μs (0 allocations: 0 bytes)
@btime @strided $B .= ($A .* $A);
# 6.216 μs (26 allocations: 3.50 KiB)

```

I’m wondering if you could describe briefly what is happening under the hood that makes this possible? I was under the impression Base broadcasting also tried to do smart things with threads (although I’m ignorant of what exactly that is tbh). What is Strided doing beyond that?

---

<div class="post-metadata">

**Author:** ![jebej](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jebej/32/1784_2.png) [@jebej](https://discourse.julialang.org/u/jebej)\
**Post date:** [May 22, 2019, 12:13am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/7 "2019-05-22T00:13:59Z")

</div>

As of yet, broadcasting does not use threads.

The relevant issue for discussion is [here](https://github.com/JuliaLang/julia/issues/19777).

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [May 22, 2019, 12:52am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/8 "2019-05-22T00:52:27Z")

</div>

Beautiful contribution! Looking forward to more speedups in the language! 🙂

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [May 22, 2019, 2:02am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/9 "2019-05-22T02:02:36Z")

</div>

Does the performance depend a lot on the architecture? I don’t get any speed up:

```julia
julia> using Strided                                                      
julia> using BenchmarkTools                                                
julia> A = randn(256,256);                                                 
julia> B = similar(A);                                                     
julia> @btime $B .= ($A .* $A);                                                      
  21.333 μs (0 allocations: 0 bytes)                              
julia> @btime @strided $B .= ($A .* $A);                                             
  21.797 μs (16 allocations: 720 bytes)              

```

for

```julia
julia> versioninfo()                                                                 
Julia Version 1.1.0                                                                  
Commit 80516ca202 (2019-01-21 21:24 UTC)                                             
Platform Info:                                                                       
  OS: Windows (x86_64-w64-mingw32)                                                   
  CPU: Intel(R) Core(TM) i7-6650U CPU @ 2.20GHz                                      
  WORD_SIZE: 64                                                                      
  LIBM: libopenlibm                                                                  
  LLVM: libLLVM-6.0.1 (ORCJIT, skylake)   

```

---

<div class="post-metadata">

**Author:** ![tshort](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tshort/32/43_2.png) [@tshort](https://discourse.julialang.org/u/tshort)\
**Post date:** [May 22, 2019, 2:17am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/10 "2019-05-22T02:17:15Z")

</div>

@PetrKryslUCSD, I think you need to be using more than one thread to see a speedup here.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [May 22, 2019, 2:26am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/11 "2019-05-22T02:26:05Z")

</div>

I dont know, it says at the top that Strided is not multithreaded…?

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [May 22, 2019, 3:51am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/12 "2019-05-22T03:51:53Z")

</div>

I get no speedup using `@strided` with 1 thread.  
I get 3x speedup using `@strided` with 6 threads.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [May 22, 2019, 4:24am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/13 "2019-05-22T04:24:20Z")

</div>

> [@juthohaegeman](#):
>
> Since these implementations are also multithreaded (using `Threads.@threads` ) there might even be a speedup if all your arrays are just plain `Array` s with monotoneously increasing strides, provided you have `Threads.nthreads() > 1` .

Doesn’t the OP say that it _is threaded_?

---

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [May 22, 2019, 6:59am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/14 "2019-05-22T06:59:44Z")

</div>

Yes, Strided.jl uses/exploits the built-in multithreading capabilities of Julia. Basically, given an array problem, say broadcasting or map, with one or more arrays involved, it will first divide the problem into a a number of disconnected subproblems/blocks (based on some heuristic as to which dimensions to slice up), equal to the number of threads, and then use a simple `@threads for` to loop over these subproblems / blocks. Normally, reduction dimensions are not sliced, as this would require temporary storage to collect the result. The only exception is a full reduction, as in that case the required temporary is small (a vector of length equal to the number of threads) and there would otherwise be no possibility to exploit more than one thread.

To see any effect, one needs to start julia with more than one thread. The only way I think this is possible is to set the environment variable `JULIA_NUM_THREADS=...` before Julia starts. I think Juno also sets this variable to a number larger than one (equal to the number of cores presumably).

Furthermore, very simple operations such as a plain `copy!` or maybe a simple addition of two arrays will likely not see any speedup, as they are indeed bandwidth limited. But if more complex operations are involved, there could be some speedup (but rarely equal to the number of threads).

---

<div class="post-metadata">

**Author:** ![juthohaegeman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juthohaegeman/32/8620_2.png) [@juthohaegeman](https://discourse.julialang.org/u/juthohaegeman)\
**Post date:** [May 22, 2019, 7:01am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/15 "2019-05-22T07:01:31Z")

</div>

> [@PetrKryslUCSD](#):
>
> I dont know, it says at the top that Strided is not multithreaded…?

Where does it say that? If you read this somewhere, this must have been a typo?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [May 22, 2019, 2:01pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/16 "2019-05-22T14:01:42Z")

</div>

My apologies, I misread.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [May 22, 2019, 3:02pm UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/17 "2019-05-22T15:02:18Z")

</div>

The computation with `@strided` hangs with more than one thread, win 10, julia 1.3.0-DEV263. Has anyone else seen this?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [June 4, 2019, 1:26am UTC](https://discourse.julialang.org/t/ann-strided-jl-a-high-level-dense-array-acceleration-package/16434/18 "2019-06-04T01:26:40Z")

</div>

Disappointing performance with WSL on Windows 10:

```julia
julia> using Strided
julia> using BenchmarkTools
julia> A = randn(256,256);
julia> B = similar(A);
julia> @btime $B .= ($A .* $A);
  22.000 μs (0 allocations: 0 bytes)
julia> @btime @strided $B .= ($A .* $A);
  32.000 μs (34 allocations: 2.66 KiB)
julia> versioninfo()
Julia Version 1.3.0-DEV.348
Commit b0086f21b1 (2019-06-03 20:19 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Core(TM) i7-6650U CPU @ 2.20GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.1 (ORCJIT, skylake)
Environment:
  JULIA_NUM_THREADS = 2

```

Excellent, works very well with Windows 10 itself:

```julia
julia> using Strided
julia> using BenchmarkTools
julia> A = randn(256,256);
julia> B = similar(A);
julia> @btime $B .= ($A .* $A);
  21.797 μs (0 allocations: 0 bytes)
julia> @btime @strided $B .= ($A .* $A);
  12.985 μs (20 allocations: 1.17 KiB)
julia> versioninfo()
Julia Version 1.1.0
Commit 80516ca202 (2019-01-21 21:24 UTC)
Platform Info:
  OS: Windows (x86_64-w64-mingw32)
  CPU: Intel(R) Core(TM) i7-6650U CPU @ 2.20GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.1 (ORCJIT, skylake)
Environment:
  JULIA_EDITOR = "C:\Users\PetrKrysl\AppData\Local\atom\app-1.37.0\atom.exe" -a
  JULIA_NUM_THREADS = 2

```

On a “real” Linux machine ;), also very nice:

```julia
julia> using Strided
julia> using BenchmarkTools
julia> A = randn(4*256,4*256);
julia> B = similar(A);
julia> @btime $B .= ($A .* $A);
  2.735 ms (0 allocations: 0 bytes)
julia> @btime @strided $B .= ($A .* $A);
  73.374 μs (30 allocations: 11.28 KiB)
julia> versioninfo()
Julia Version 1.1.0
Commit 80516ca202 (2019-01-21 21:24 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: AMD Opteron(tm) Processor 6380
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-6.0.1 (ORCJIT, bdver1)
Environment:
  JULIA_NUM_THREADS = 64

```
