# Improving parallel loop performance by using \`numpy\` for allocations

**URL:** <https://discourse.julialang.org/t/improving-parallel-loop-performance-by-using-numpy-for-allocations/122420>\
**Category:** Performance\
**Created:** [November 8, 2024, 2:04pm UTC](https://discourse.julialang.org/t/improving-parallel-loop-performance-by-using-numpy-for-allocations/122420 "2024-11-08T14:04:08Z")\
**Posts on this page:** 1\
**Showing post:** 3

<div class="post-metadata">

**Author:** ![artemsolod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/artemsolod/32/20704_2.png) [@artemsolod](https://discourse.julialang.org/u/artemsolod)\
**Post date:** [November 10, 2024, 4:11pm UTC](https://discourse.julialang.org/t/improving-parallel-loop-performance-by-using-numpy-for-allocations/122420/3 "2024-11-10T16:11:07Z")

</div>

Speed gains are visible in single-threaded workloads also. Seemingly they are not purely due to `GC` pauses - the timings are similar with `GC` disabled. Here is an example of `copy and sum a vector`. Having numpy managing memory gives a good speedup on both x86 and Mac, and for different (big-ish) vector size.  
Timings:

```julia
julia +release --project=. -t 1 copyadd_time.jl 
  0.256724 seconds (64.30 k allocations: 79.529 MiB, 53.59% gc time, 17.37% compilation time) # compile
  2.745950 seconds (4.35 M allocations: 218.689 MiB, 3.69% gc time, 97.00% compilation time) # compile
  0.054762 seconds (46 allocations: 1.367 KiB) # numpy/pyarray
  0.074416 seconds (4 allocations: 76.294 MiB) # native julia

```

Code:

```julia
ENV["JULIA_CONDAPKG_BACKEND"] = "Null" # use system-wide python installation
# otherwise install numpy for this PythonCall environment:
# ] add CondaPkg; using CondaPkg; ] conda add numpy
using PythonCall
using Random
np = pyimport("numpy")
Random.seed!(42)

function copy_jl(arr)
    sum(copy(arr))
end

function copy_np(arr)
    pymem = np.empty(length(arr))
    pyarr = PyArray(pymem)
    pyarr .= arr
    ans = sum(pyarr)
    return ans
end

arr = rand(10_000_000)

@time copy_jl(arr)
@time copy_np(arr)
GC.gc()
GC.enable(false) # doesn't matter

@time copy_np(arr)
@time copy_jl(arr)

```

---

_[View the full topic](https://discourse.julialang.org/t/improving-parallel-loop-performance-by-using-numpy-for-allocations/122420)._
