# What is the best way to re-use a temporary vector

**URL:** <https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519>\
**Category:** Performance\
**Tags:** performance, memory-allocation\
**Created:** [October 21, 2024, 2:24am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519 "2024-10-21T02:24:06Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![HMegh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hmegh/32/216684_2.png) [@HMegh](https://discourse.julialang.org/u/HMegh)\
**Post date:** [October 21, 2024, 2:24am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/1 "2024-10-21T02:24:06Z")

</div>

Hi,

I have the following situation: I have a function `f(x::Float64)` which creates a temporary array `A::Vector{Float64}`, then finally it returns a `Float64`. I typically call this function millions of times, so I am wondering if there is a way to allocate a single vector `A` and pass it to each function call efficiently.

* * *

**MWE :** The `dnPl` function from [LegendrePolynomials.jl](https://github.com/jishnub/LegendrePolynomials.jl) uses a _cache_ to calculate the n-th derivative of a Legendre polynomial. Here’s an example:

```julia
using BenchmarkTools, LegendrePolynomials

@btime dnPl(.5,10,4)
# 61.842 ns (2 allocations: 112 bytes)

cache=zeros(10) #temporary cache
@btime dnPl(.5,10,4,$cache)
# 49.053 ns (0 allocations: 0 bytes)

```

In my code, I need to evaluate `dnPl` for many inputs. For example

```julia
function without_cache(n) 
    A=zeros(n,10) 
    for i=1:n,j=1:10
        x=-1+2i/n
        A[i,j]=dnPl(x,j,4)
    end 
    A 
end

@btime without_cache(30)
# 22.913 μs (693 allocations: 25.42 KiB)

```

The number of allocations seem to be consistent with the previous example: (2 allocations per `dnPl` call) times 300, plus some overhead. So, it tried to create a local `cache`. However, I still see too many allocations

```julia
function with_cache(n) 
    A=zeros(n,10) 
    cache=zeros(10)
    for i=1:n,j=1:10
        x=-1+2i/n
        A[i,j]=dnPl(x,j,4,cache)
    end 
    A 
end

@btime with_cache(30)
# 19.767 μs (275 allocations: 6.81 KiB)

```

Do you have any ideas on how to re-use the “cache” without having these allocations?

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [October 21, 2024, 2:43am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/2 "2024-10-21T02:43:44Z")

</div>

Use a callable object? (For example: [FinEtoolsFlexStructures.jl/src/FEMMShellT3FFModule.jl at 9b70a29592010aa9edeabc8f47aa348ebadc223b · PetrKryslUCSD/FinEtoolsFlexStructures.jl · GitHub](https://github.com/PetrKryslUCSD/FinEtoolsFlexStructures.jl/blob/9b70a29592010aa9edeabc8f47aa348ebadc223b/src/FEMMShellT3FFModule.jl#L377))

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [October 21, 2024, 3:17am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/3 "2024-10-21T03:17:55Z")

</div>

Actually, I think `dnPl` allocates.

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [October 21, 2024, 3:37am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/4 "2024-10-21T03:37:56Z")

</div>

Unrelated to your question about allocations, but you probably want to flip the order of your loops (i,j) to access A according to column major order (to avoid cache misses).

---

<div class="post-metadata">

**Author:** ![Satvik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/satvik/32/20486_2.png) [@Satvik](https://discourse.julialang.org/u/Satvik)\
**Post date:** [October 21, 2024, 3:45am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/5 "2024-10-21T03:45:40Z")

</div>

You could use Bumper.jl

---

<div class="post-metadata">

**Author:** ![HMegh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hmegh/32/216684_2.png) [@HMegh](https://discourse.julialang.org/u/HMegh)\
**Post date:** [October 21, 2024, 3:00pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/6 "2024-10-21T15:00:04Z")

</div>

I might be missing something here but even `Bumper.jl` seem to produce a similar number of allocations (I could be misusing though)

```julia
function with_bumper(n)
    A=zeros(n,10) 
    @no_escape begin 
        cache=@alloc(Float64,10)
        for j=1:10, i=1:n
            x=-1+2i/n
            A[i,j]=dnPl(x,j,4,cache)
        end 
    end #no_escape
    A 
end

@btime with_bumper(30)
# 18.692 μs (273 allocations: 6.67 KiB)

```

It could be the case that `dnPl` is allocating as @PetrKryslUCSD mentioned. However, I did not see any allocation (see my original post).

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [October 21, 2024, 3:15pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/7 "2024-10-21T15:15:24Z")

</div>

It might be that it allocates for some inputs.

---

<div class="post-metadata">

**Author:** ![HMegh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hmegh/32/216684_2.png) [@HMegh](https://discourse.julialang.org/u/HMegh)\
**Post date:** [October 21, 2024, 3:27pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/8 "2024-10-21T15:27:09Z")

</div>

Thanks, I went ahead and attempted that: There are no allocations when I am simply calling the object but the problem happens when I have the cache outside of a loop:

```julia
mutable struct my_dnPl
    x::Float64
    n::Int64
    k::Int64 
    cache::Vector{Float64}
end
function (p::my_dnPl)() 
    return dnPl(p.x,p.n,p.k,p.cache) 
end

cache=zeros(10)
@btime my_dnPl(.4,10,3,$cache)()
# 63.603 ns (0 allocations: 0 bytes

)
function with_callable_obj(n)
    A=zeros(n,10) 
    cache=zeros(10)
    for i=1:n,j=1:10
        x=-1+2i/n
        A[i,j]=my_dnPl(x,j,4,cache)()
    end 
    A 
end

@btime with_callable_obj(30)
 # 21.728 μs (575 allocations: 20.88 KiB)

```

When `struct` is used instead of `mutable struct` I still see `275 allocations`. So, maybe `dnPl` is allocating somewhere that I can’t see in the code.

Having said that, I have found a workaround using another function that the same package that solves this problem, which I will include below in case someone runs into the same issue

```julia
function with_collect(n) 
    A=zeros(n,10) 
    cache=zeros(10+1) #collectdnPl also returns P_0
    for i=1:n 
        x=-1+2i/n
        collectdnPl!(cache,x,lmax=10,n=4)
        A[i,:]=@views cache[2:end] 

    end
    A
end

@btime with_collect(30) 
# 2.204 μs (5 allocations: 2.59 KiB)

```

---

<div class="post-metadata">

**Author:** ![Satvik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/satvik/32/20486_2.png) [@Satvik](https://discourse.julialang.org/u/Satvik)\
**Post date:** [October 21, 2024, 3:29pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/9 "2024-10-21T15:29:05Z")

</div>

It does seem like `dnPl` allocates, though not very much or often:

```julia
julia> function with_cache(n)
           A=zeros(n,10)
           cache=zeros(10)
           for i=1:n,j=1:10
               @timeit "x" x=-1+2i/n
               @timeit "A" A[i,j]=dnPl(x,j,4,cache)
           end
           A
       end
with_cache (generic function with 1 method)

julia> @btime with_cache(30);
  83.917 μs (272 allocations: 6.98 KiB)

julia> print_timer()
 ────────────────────────────────────────────────────────────────────
                            Time Allocations
                   ─────────────────────── ────────────────────────
 Tot / % measured: 14.3s / 23.9% 475MiB / 57.8%

 Section ncalls time %tot avg alloc %tot avg
 ────────────────────────────────────────────────────────────────────
 A 20.0M 2.97s 86.8% 149ns 274MiB 100.0% 14.4B
 x 20.0M 454ms 13.2% 22.7ns 0.00B 0.0% 0.00B
 ────────────────────────────────────────────────────────────────────

```

---

<div class="post-metadata">

**Author:** ![StevenSiew](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevensiew/32/218393_2.png) [@StevenSiew](https://discourse.julialang.org/u/StevenSiew)\
**Post date:** [October 21, 2024, 9:09pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/10 "2024-10-21T21:09:19Z")

</div>

How to speed up loops in Matrix

```julia
Slow for big matrices

# Slow way, first loop in col then loop in row
function incmatrixSLOW(a)
    (m,n) = size(a)
    for row=1:m , col=1:n
        a[row,col] += 1
    end
end

Fast for big matrices

# fast way, first loop in row then loop in col
function incmatrixFAST(a)
    (m,n) = size(a)
    for col=1:n , row=1:m 
        a[row,col] += 1
    end
end

Fast by getting julia to do it for you

function incmatrixFAST(a)
    (m,n) = size(a)
    for k = eachindex(a) 
        a[k] += 1
    end
end

```

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [October 22, 2024, 10:52am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/11 "2024-10-22T10:52:39Z")

</div>

You can fix that via

```julia
julia> @eval LegendrePolynomials function doublefactorial(::Type{T}, n) where T
           p = convert(T, one(n))
           for i in n:-2:1
               p *= convert(T, i)
           end
           convert(T, p)
       end

```

in order to replace the original

```julia
function doublefactorial(T, n)
           p = convert(T, one(n))
           for i in n:-2:1
               p *= convert(T, i)
           end
           convert(T, p)
       end

```

Somebody might need to do a PR on that for LegendrePolynomials.jl.

However, I would much more interested in _why_ this goes wrong here and getting it fixed in the compiler!

This is clearly a failure of constant-propagation / inlining / specialization / type-inference. You don’t see it by benchmarking dnPl in isolation because this has different state of inlining / const-prop heuristics.

The bad pattern – using `T` as an argument when you really mean `::Type{T}` – is very commonly used. It would be better if we could make it reliably do the right thing.

PS. I looked at

```julia
julia> import Profile
julia> with_cache(30); Profile.Allocs.clear();Profile.Allocs.@profile with_cache(30); p = Profile.Allocs.fetch(); Profile.print(p)
Overhead ╎ [+additional indent] Count File:Line; Function
=========================================================
   ╎432 @Base/client.jl:541; _start()
   ╎ 432 @Base/client.jl:567; repl_main
   ╎ 432 @Base/client.jl:430; run_main_repl(interactive::Bool, quiet::Bool, banner::Symbol, history_file::Bool, color_set::Bool)
   ╎ 432 @Base/essentials.jl:1052; invokelatest
   ╎ 432 @Base/essentials.jl:1055; #invokelatest#2
   ╎ 432 @Base/client.jl:446; (::Base.var"#1139#1141"{Bool, Symbol, Bool})(REPL::Module)
   ╎ ╎ 432 @REPL/src/REPL.jl:469; run_repl(repl::REPL.AbstractREPL, consumer::Any)
   ╎ ╎ 432 @REPL/src/REPL.jl:483; run_repl(repl::REPL.AbstractREPL, consumer::Any; backend_on_current_task::Bool, backend::Any)
   ╎ ╎ 432 @REPL/src/REPL.jl:324; kwcall(::NamedTuple, ::typeof(REPL.start_repl_backend), backend::REPL.REPLBackend, consumer::Any)
   ╎ ╎ 432 @REPL/src/REPL.jl:327; start_repl_backend(backend::REPL.REPLBackend, consumer::Any; get_module::Function)
   ╎ ╎ 432 @REPL/src/REPL.jl:342; repl_backend_loop(backend::REPL.REPLBackend, get_module::Function)
   ╎ ╎ ╎ 432 @REPL/src/REPL.jl:245; eval_user_input(ast::Any, backend::REPL.REPLBackend, mod::Module)
   ╎ ╎ ╎ 432 @Base/boot.jl:430; eval
   ╎ ╎ ╎ 432 REPL[3]:6; with_cache(n::Int64)
   ╎ ╎ ╎ 432 @LegendrePolynomials/src/LegendrePolynomials.jl:352; dnPl
   ╎ ╎ ╎ 432 @LegendrePolynomials/src/LegendrePolynomials.jl:284; _unsafednPl!
 64╎ ╎ ╎ ╎ 64 @LegendrePolynomials/src/LegendrePolynomials.jl:236; doublefactorial(T::Type, n::Int64)
368╎ ╎ ╎ ╎ 368 @LegendrePolynomials/src/LegendrePolynomials.jl:238; doublefactorial(T::Type, n::Int64)
Total snapshots: 27
Total bytes: 432

```

From there you can take a look at the relevant source code and you immediately see the code-smell (proper spelling has `::Type{T}` as argument, also that should be `@inline`).

On the other hand, seeing the code in isolation, I would not have expected that to actually create trouble! Especially given that it works fine when you benchmark `dnPl` in isolation.

---

<div class="post-metadata">

**Author:** ![nsajko](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nsajko/32/221187_2.png) [@nsajko](https://discourse.julialang.org/u/nsajko)\
**Post date:** [October 22, 2024, 11:26am UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/12 "2024-10-22T11:26:14Z")

</div>

> [@foobar\_lv2](#):
>
> The bad pattern – using `T` as an argument when you really mean `::Type{T}` – is very commonly used. It would be better if we could make it reliably do the right thing.

There is no _the_ right thing, one has to be wary of both over- and under- specialization while programming. In fact, what I’d like to see is more support for explicit control over what gets specialized. Relevant issue:

> <https://github.com/JuliaLang/julia/issues/11339>
>
> aka "Generalize the ANY mechanism" aka "Completely destroy Jeff's life forever"
> …(continuing the heresy in https://github.com/JuliaLang/julia/issues/4774#issuecomment-103286221 in a direction that might be more profitable, but much harder)
> 
> What if we could encode an Array's size among its type parameters, but not pay the compilation price of generating specialized code for each different size?
> 
> Here's a candidate set of rules (EDITED):
> 1. In the type declaration, you specify the default specialization behavior for the parameters: \`type Bar{T,K\*}\` means that, by default, functions are specialized on \`T\` but not on \`K\`. Array might be defined as \`type Array{T,N,SZ\*\<:NTuple{N,Int}}\` (in a future world of triangular dispatch).
> 2. In a function signature, if a type's parameter isn't represented as a "TypeVar", it uses the default behavior. So \`foo(A::Array)\` specializes the element type and the dimensionality, but not the size.
> 3. Specifying a parameter overrides the default behavior. \`foo{T,K}(b::Bar{T,K})\` would specialize on \`K\`, for example. Conversely, putting \\\* after a type parameter causes the parameter to be used in dispatch (subtying and intersection) but prevents specialization. This is much like the ANY hack now. \`matmul{S,T,m\*,k\*,n\*}(A::Array{S,2,(m,k)}, B::Array{T,2,(k,n)}\` would perform the size check at compile time but not specialize on those arguments.
> 
> I'm aware that this is more than a little crazy and has some serious downsides (among others, it feels like performance could get finicky). But to go even further out on a limb, one might additionally have \`matmul{S,T,m\<=4,k\<=4,n\<=4}(...)\` to generate specialized algorithms (note the absence of \*s) for small sizes and use the unspecialized fallbacks for larger sizes. This way one wouldn't need a specialized \`FixedSizeArray\` type, and could just implement such algorithms on \`Array\`.
> 
> A major downside to this proposal is that calls in functions that lack complete specialization might have to be expanded like this:
> 
> \`\`\` jl
> if size(A,1) \<= 4 && size(A,2) \<= 4 && size(B,2) \<= 4
> C = matmul\_specialized(A,B)
> else
> C = matmul\_unspecialized(A,B)
> end
> \`\`\`

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [October 22, 2024, 12:05pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/13 "2024-10-22T12:05:20Z")

</div>

> [@nsajko](#):
>
> There is no _the_ right thing, one has to be wary of both over- and under- specialization while programming.

Yes. And in this specific case, it is very obvious that the LegendrePolynomials.jl author meant `::Type{T}`, i.e. wanted to specialize.

But inlining + const-prop does the right specialization despite their mistake in many contexts like e.g. `@btime dPnl($x, $n, $ell, $cache)`. So we end up with a situation where reasonable-ish testing does not uncover a serious performance bug in a library ☹

I’m not sure how to approach this. Maybe we need a command-line switch / build-time switch for debug/testing that only gives you forced specializations? Such that context can never hide such a bug?

The specialization rules are full of weirdness already, like the magic rule that “if the function is called, then it is specialized”. Sure.

Why not “if a function body contains a loop, then it is specialized”?

Or why not get rid of the magic “don’t specialize on \<:Function or \<:Type” rule completely?

I’m not saying that always specializing is the right thing. It’s more a question of which bugs happen for which people and how hard they are to spot.

My guess is that the people who would need many more `@nospecialize` annotations are the ones who are more aware of [Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/#Be-aware-of-when-Julia-avoids-specializing) than the people who today need to force specializations in their code.

---

<div class="post-metadata">

**Author:** ![HMegh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hmegh/32/216684_2.png) [@HMegh](https://discourse.julialang.org/u/HMegh)\
**Post date:** [October 22, 2024, 12:50pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/14 "2024-10-22T12:50:21Z")

</div>

This fixes the issue, thank you. I will run a few tests, then I will submit a PR. I have a question though: What does `@eval` do here?

---

<div class="post-metadata">

**Author:** ![abraemer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/abraemer/32/51403_2.png) [@abraemer](https://discourse.julialang.org/u/abraemer)\
**Post date:** [October 22, 2024, 1:10pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/15 "2024-10-22T13:10:37Z")

</div>

It evaluates the code block inside the module `LegendrePolynomials`. Basically exactly as if you had typed it inside the package’s code. Thus it overwrites the original definition.

I am not sure though what the difference is to just overwriting the method

```julia-repl
julia> function LegendrePolynomials.doublefactorial(::Type{T}, n) where T
           p = convert(T, one(n))
           for i in n:-2:1
               p *= convert(T, i)
           end
           convert(T, p)
       end

```

---

<div class="post-metadata">

**Author:** ![mike.ingold](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mike.ingold/32/203749_2.png) [@mike.ingold](https://discourse.julialang.org/u/mike.ingold)\
**Post date:** [October 22, 2024, 2:51pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/16 "2024-10-22T14:51:13Z")

</div>

> [@HMegh](#):
>
> I have the following situation: I have a function `f(x::Float64)` which creates a temporary array `A::Vector{Float64}`, then finally it returns a `Float64`. I typically call this function millions of times, so I am wondering if there is a way to allocate a single vector `A` and pass it to each function call efficiently.

> [@HMegh](#):
>
> In my code, I need to evaluate `dnPl` for many inputs. For example
> 
> ```julia
> function without_cache(n) 
> A=zeros(n,10) 
> for i=1:n,j=1:10
> x=-1+2i/n
> A[i,j]=dnPl(x,j,4)
> end 
> A 
> end
> 
> ```

I’m not sure if these are related statements since this function returns an array of size `(n, 10)` instead of a scalar. However, whenever I see an array being allocated to store highly-structured contents, my mind immediately goes to using Iterators to eliminate allocations. For the function above, this might look something like

```julia
function with_iterators(n) 
    i = 1:n
    j = 1:10
    ij = Iterators.product(i, j)
    Iterators.map((i,j) -> dnPl(-1 + 2i/n, j, 4), ij)
end

```

which should produce an iterator that lazily evaluates the `dnPl` call at `getindex` time instead of allocating memory to store all of the results.

If you’re going to immediately reduce the array anyway, like the function you mentioned at the top where ultimately `f = x::Float64 -> y::Float64`, you can then perform the reduction without ever allocating in the first place, e.g. `sum(f, with_iterators(n))`, `reduce(f, with_iterators(n))`, etc.

---

<div class="post-metadata">

**Author:** ![HMegh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hmegh/32/216684_2.png) [@HMegh](https://discourse.julialang.org/u/HMegh)\
**Post date:** [October 22, 2024, 4:24pm UTC](https://discourse.julialang.org/t/what-is-the-best-way-to-re-use-a-temporary-vector/121519/17 "2024-10-22T16:24:11Z")

</div>

Thanks, using `Iterators` instead of allocating `A` is definitely better. However, the issue here apparently has to do with `dnPl` and the way a certain function is declared.
