# Y.A.t.Q : Yet Another @threads Question

**URL:** <https://discourse.julialang.org/t/y-a-t-q-yet-another-threads-question/71541>\
**Category:** General Usage\
**Tags:** parallel, multithreading, threads\
**Created:** [November 16, 2021, 7:56am UTC](https://discourse.julialang.org/t/y-a-t-q-yet-another-threads-question/71541 "2021-11-16T07:56:02Z")\
**Posts on this page:** 1\
**Showing post:** 13

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [November 16, 2021, 3:48pm UTC](https://discourse.julialang.org/t/y-a-t-q-yet-another-threads-question/71541/13 "2021-11-16T15:48:37Z")

</div>

If you care for performance of array traversal code, it’s crucial to consider memory layout. In Julia, arrays are column-major by default, and switching to row-major gives almost a 5-fold speedup for me. This is without paralellization:

```julia
X_sub_r = PermutedDimsArray(permutedims(X_sub), (2, 1))
X_r = PermutedDimsArray(permutedims(X), (2, 1))
sequenceCosine(X_sub, X) # 3.6 s
sequenceCosine(X_sub_r, X_r) # 0.72 s

```

Writing this function in terms of `map` can make it absolutely trivial to parallelize:

```julia
# sequential - has the same performance as sequenceCosine, 0.7 s for row-major arrays
function sequenceCosine3(X_sub, X)
	map(Iterators.product(eachrow(X_sub), eachrow(X))) do (i, j)
    	cosine_dist(i, j)
	end
end

# parallel - 0.26 s for me
function parallelCosine3(X_sub, X)
	ThreadsX.map(Iterators.product(eachrow(X_sub), eachrow(X))) do (i, j)
    	cosine_dist(i, j)
	end
end

```

---

_[View the full topic](https://discourse.julialang.org/t/y-a-t-q-yet-another-threads-question/71541)._
