# Array of functions - is there a way to avoid allocations performance penalty?

**URL:** <https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471>\
**Category:** Performance\
**Tags:** memory-allocation, arrays\
**Created:** [May 22, 2019, 11:29am UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471 "2019-05-22T11:29:51Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![donboyd5](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/donboyd5/32/8562_2.png) [@donboyd5](https://discourse.julialang.org/u/donboyd5)\
**Post date:** [May 22, 2019, 11:29am UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/1 "2019-05-22T11:29:52Z")

</div>

_Edited version of my initial post based upon the response from @Raf._

After I wrote this post (1) I got advice on how to write a better post. I have taken that advice below, and (2) I actually found the answer to my post by looking at [Understanding the allocation behavior on arrays vs scalars - #2 by stevengj](https://discourse.julialang.org/t/understanding-the-allocation-behavior-on-arrays-vs-scalars/23798/2) . Thus, I am editing the post to include the answer, in case others have the same question.

I want to maximize performance when I call functions that are array elements. My underlying motivation is that I plan to write income tax simulation models (tax calculators) for tax systems in ~41 states. The models will run on microdata files that represent taxpayers in the states. I want the models to be as parameter-driven as possible, where functions that do the necessary tax calculations are stored in arrays. The actual functions, the order in which they are processed, and their arguments, will vary from state to state, and from policy scenario (potential tax law) to policy scenario, defined in json files. It is a variant of and inspired by work being done [here](https://github.com/PSLmodels/Tax-Calculator).

I was concerned because when I tested the idea of putting functions into arrays, performance degraded substantially. I have since discovered from the post linked above that this problem can be solved by using StaticArrays.

The remainder of this post shows the problem, and the solution.

**The problem**  
The example below first defines consts for two functions (sin, cos) separately, and then defines a const array of these two functions. Then I call each of the two functions 10^6 times, two different ways: (1) calling the consts, and (2) calling the array elements. (I have constructed the example this way to isolate the issue, so that we can see - I believe - that it results solely from using an array of functions.)

The first approach has 13.5k memory allocations on the first run and only 6 allocations on the second run (6 allocations for @time I believe).

The second approach (array elements) has 6 M allocations on each call (and takes far more time when scaled up) and so is much less efficient. Because this is an approach I would like to be able to use, my initial question was, is there a way to make it more efficient?

```julia
const fvar1 = sin
const fvar2 = cos
const fvec = [sin, cos]

function fun_vars(n::Int64)
    t = 0.0
    for i = 1:n
        t += fvar1(i)
        t += fvar2(i)
    end
    return t
end

function fun_vec(n::Int64)
    t = 0.0
    for i = 1:n
        t += fvec[1](i)
        t += fvec[2](i)
    end
    return t
end

@time fun_vars(10^6) # 13.5k allocations
@time fun_vars(10^6) # 6 allocations on 2nd time

@time fun_vec(10^6) # 6 M allocations
@time fun_vec(10^6) # 6 M allocations on 2nd time

```

**The solution**  
It turns out there is a way to make this more efficient. A solution - perhaps the best solution, I do not know - is to use StaticArrays, as shown below.

```julia
# Pkg.add("StaticArrays") 
using StaticArrays

const fvec2 = SVector(sin, cos)

function fun_vec2(n::Int64)
    t = 0.0
    for i = 1:n
        t += fvec2[1](i)
        t += fvec2[2](i)
    end
    return t
end

@time fun_vec2(10^6) # 16.6k allocations
@time fun_vec2(10^6) # 6 allocations on 2nd time

```

I hope this is helpful to someone.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [May 22, 2019, 11:36am UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/2 "2019-05-22T11:36:40Z")

</div>

Welcome!

Please put your code in a code block so it’s well formatted - so we can read it.

> [@Please read: make it easier to help you](https://discourse.julialang.org/t/psa-make-it-easier-to-help-you/14757):
>
> Welcome to the Julia Discourse! We are enthusiastic about helping Julia programmers, both beginner and experienced. This public service announcement (PSA) outlines best practices when asking for help. Following these points makes it easier for us to help you and more likely you’ll get a prompt, useful answer. Keywords are highlighted to make it easier to refer to specific points. Choose a descriptive title that captures the key part of your question, eg “plots with multiple axes” instead of …

With a question like this, you could also describe what you are actually trying to do: the wider purpose of this code. Helping you make a better array of functions is not really helping you in the long term - you probably should be doing something else entirely. A little background will help people work out what that something is.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [May 22, 2019, 12:13pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/3 "2019-05-22T12:13:40Z")

</div>

Static arrays could be better, because they are really tuples underneath. You could instead just use a tuple.

Personally I would use a tuple of structs that are all passed to the same function, but they have their own methods defined for that function. That means you can attach additional parameters for each function if you need them, making it a lot more extensible. But it could also be overkill.

---

<div class="post-metadata">

**Author:** ![donboyd5](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/donboyd5/32/8562_2.png) [@donboyd5](https://discourse.julialang.org/u/donboyd5)\
**Post date:** [May 22, 2019, 12:16pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/4 "2019-05-22T12:16:23Z")

</div>

Many thanks. I will learn about this.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [May 22, 2019, 1:29pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/5 "2019-05-22T13:29:24Z")

</div>

> [@Raf](#):
>
> Personally I would use a tuple of structs that are all passed to the same function, but they have their own methods defined for that function. That means you can attach additional parameters for each function if you need them, making it a lot more extensible. But it could also be overkill.

Currently, I’m working on a similar issue as @donboyd5 where I need an array of functions available.

In my case, I will be having anywhere from 1 to 64 functions in the array. In many cases, I will only need a few specific values from the whole array and sometimes I will need to calculate from all the functions in the array simultaneously.

For 1-64 functions in an array, with the need for computing individual values in some cases, and all values in other cases, would it be better to make a single function call, or individual calls? Hmm

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [May 22, 2019, 1:37pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/6 "2019-05-22T13:37:12Z")

</div>

I don’t mean a single function call, but calling some `run_it` function for each struct in the tuple, triggering the appropriate method that matches the struct. The advantage is having self contained units that can hold additional parameters or preallocated arrays where required (if you need to run them millions of times).

I do this to build add-hoc composite models in my Celluar.jl and Dispersal.jl , GrowthRates.jl packages. But they are short (max 10 in a chain), and many need parameters - so I use a chain of structs that are passed to the same function in sequence. YMMV

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [May 22, 2019, 1:38pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/7 "2019-05-22T13:38:29Z")

</div>

> [@Raf](#):
>
> I don’t mean a single function call, but calling some `run_it` function for each struct in the tuple, triggering the appropriate method that matches the struct. The advantage is having self contained units that can hold additional parameters or preallocated arrays where required (if you need to run them millions of times).

Could you provide a minimal example?

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [May 22, 2019, 1:42pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/8 "2019-05-22T13:42:02Z")

</div>

[GitHub - yuyichao/FunctionWrappers.jl](https://github.com/yuyichao/FunctionWrappers.jl) was created for this purpose (see [[RFC] Type stable wrapper for arbitrary callables · Issue #13984 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/13984)). I am however not sure how well it works on the latest julia release.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [May 22, 2019, 1:54pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/9 "2019-05-22T13:54:02Z")

</div>

It seems you are not using the fact that the vector is a vector. A tuple should work equally well for your example?

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [May 22, 2019, 1:54pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/10 "2019-05-22T13:54:49Z")

</div>

This is only worth doing if your functions are suitably complicated, and need parameters or preallocated arrays occasionally, which is a lot of the time for me. Eg. one such function needs a precalculated dispersal kernel attached, and identical sized array for doing blas multiplications.

```julia
# Interface
function dosomething end

# Functions
struct Square end
dosomething(t::Square, x) = x^2

struct ParameterMult{P}
    param::P
end
dosomething(t::ParameterMult, x) = t.param * x

runsim(xs, otherdata) = 
       # apply recursivly (ie type stable for a tuple) with list of structs and data, 
       # but basically like:
       dosomthing.(xs, Ref(otherdata))
end

runsim((Square(), ParameterMult(2)))

```

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [May 22, 2019, 2:23pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/11 "2019-05-22T14:23:52Z")

</div>

> [@kristoffer.carlsson](#):
>
> I am however not sure how well it works on the latest julia release.

I just wanted to take another opportunity to plug FunctionWrappers being in the stdlib or even `Base`. There aren’t that many things that you just flat out _can’t_ do in Julia, but efficient containerized functions is one of them without FunctionWrappers. Additionally, FunctionWrappers is a small amount of code, but very few people have the knowledge to implement it.

I wish I could offer help and just put in a PR for this, but it’s way over my head.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [May 22, 2019, 2:31pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/12 "2019-05-22T14:31:43Z")

</div>

Also I forgot to add, if you want this to be type stable and as fast as possible, you might need to use recursion instead of a for loop. You are looping over a tuple/SVector of different types, so it wont compile the loop to be type-stable. Unless its only a few different functions, in which case small Union optimisations might help.

Unless you are actually calling every function in the tuple manually by index as in the example, that could be type stable too.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [May 22, 2019, 3:03pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/13 "2019-05-22T15:03:55Z")

</div>

Could someone explain how exactly `FunctionWrappers` works? Is `FunctionWrapper{A,Tuple{B...}} where {A,B,...}` supposed to specify the type of the inputs and outputs with `A,B,...`?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [May 22, 2019, 3:12pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/14 "2019-05-22T15:12:43Z")

</div>

> [@chakravala](#):
>
> Could someone explain how exactly `FunctionWrappers` works? Is `FunctionWrapper{A,Tuple{B...}} where {A,B,...}` supposed to specify the type of the inputs and outputs with `A,B,...` ?

Yes, it’s a type parameterized by the input output types, so that way it can directly do the call in a way that’s type stable, but directly with that pointer and making it use that specific method (I’m not sure it can do dispatch). We’ve demonstrated that it is essentially no overhead on DiffEq

> <https://github.com/SciML/DiffEqBase.jl/pull/208>
>
> \`\`\`julia
> using OrdinaryDiffEq
> 
> \# First time 
> 
> function lorenz(du,u,p,t)
> …du\[1\] = 10.0(u\[2\]-u\[1\])
> du\[2\] = u\[1\]\*(28.0-u\[3\]) - u\[2\]
> du\[3\] = u\[1\]\*u\[2\] - (8/3)\*u\[3\]
> end
> u0 = \[1.0;0.0;0.0\]
> tspan = (0.0,100.0)
> prob = ODEProblem{true,false}(lorenz,u0,tspan)
> @time sol = solve(prob,Tsit5())
> 
> \#6.820159 seconds (16.06 M allocations: 800.508 MiB, 8.44% gc time)
> 
> \# Enable No-Recompile Mode
> 
> function lorenz2(du,u,p,t)
> du\[1\] = 10.0(u\[2\]-u\[1\])
> du\[2\] = u\[1\]\*(28.0-u\[3\]) - u\[2\]
> du\[3\] = u\[1\]\*u\[2\] - (8/3)\*u\[3\]
> end
> prob = ODEProblem{true,false}(lorenz2,u0,tspan)
> @time sol = solve(prob,Tsit5())
> 
> \#0.004248 seconds (12.96 k allocations: 1.375 MiB)
> 
> using BenchmarkTools
> 
> @btime sol = solve(prob,Tsit5())
> 
> \# 1.146 ms (12961 allocations: 1.37 MiB)
> 
> \# Disable No-Recompile Mode
> 
> function lorenz3(du,u,p,t)
> du\[1\] = 10.0(u\[2\]-u\[1\])
> du\[2\] = u\[1\]\*(28.0-u\[3\]) - u\[2\]
> du\[3\] = u\[1\]\*u\[2\] - (8/3)\*u\[3\]
> end
> prob = ODEProblem{true}(lorenz3,u0,tspan)
> @time sol = solve(prob,Tsit5())
> 
> \#1.839436 seconds (3.97 M allocations: 190.402 MiB, 4.78% gc time)
> 
> @btime sol = solve(prob,Tsit5())
> 
> \# 1.078 ms (12963 allocations: 1.37 MiB)
> \`\`\`
> 
> @YingboMa . 
> 
> Note that this isn't safe to use without \`remake\` making sure types in the FunctionWrapper change whenever changing the types, which is important when doing things like AD. But this is what would be required for static compilation to be done completely.

but of course has some interesting constraints that you have to deal with.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [May 22, 2019, 3:37pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/15 "2019-05-22T15:37:15Z")

</div>

In my case, it is needed for encoding the metric at a tangent space of a Riemannian manifold in differential geometry

> [@\[ANN\] DirectSum.jl, AbstractTensors.jl : dispatch on VectorBundle \<: Manifold{n}](https://discourse.julialang.org/t/ann-directsum-jl-abstracttensors-jl-dispatch-on-vectorspace/20870/4):
>
> The DirectSum package has a tangent bundle encoding based on the following: Let T^dV\in\text{Vect}\_\mathbb K be a VectorBundle\<:Manifold of rank n, T^dV = (n,p,g,d),\, n \in\mathbb N,\, g :V\times V\rightarrow\mathbb K, \, d \in\mathbb Z ::VectorBundle{n,p,g,d} where {n,p,g,d} byte-encoded p specifies the null-basis from the projective split of V, g is a bilinear form that specifies the metric of the space, d is an integer specifying the order of the tangent bundle. \bigoplus T^{d…

The Hodge star requires evaluating a volume element of a submanifold using the metric at a particular point, which is currently considered constant and will be generalized to a `FunctionWrapper` next.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [May 22, 2019, 4:04pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/16 "2019-05-22T16:04:45Z")

</div>

Is it possible to access the `FunctionWrapper` with `getindex` also? In my case, indexing is needed.

---

<div class="post-metadata">

**Author:** ![donboyd5](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/donboyd5/32/8562_2.png) [@donboyd5](https://discourse.julialang.org/u/donboyd5)\
**Post date:** [May 22, 2019, 8:08pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/17 "2019-05-22T20:08:34Z")

</div>

Thank you all for comments that were fast and valuable.

I realize, belatedly, that the “solution” I posted did not fully solve my problem, so here is an update that uses what I understand of what’s been written above and isolates my problem more precisely.

As two responders noted, all I need is a tuple of functions in my case, not an array.

However, when I call the functions by referencing each individual tuple member by its index, I allocate essentially no memory, but when I loop through the members, I allocate a lot of memory. The example below illustrates this.

I would like the benefits of looping through members (not needing to know the number of members, and not needing to write each function call out separately) but I would like the memory-allocation and speed benefits of not looping (i.e., essentially no memory allocation). I would much appreciate knowing whether it is possible to loop through the tuple of functions without the memory allocation penalty you see below. Thank you.

```julia

function test_elements(n::Int64)
    funs = (sin, cos)
    t = 0.0
    for i = 1:n
        t += funs[1](i)
        t += funs[2](i)
    end
    return t
end

function test_loop(n::Int64)
    funs = (sin, cos)
    t = 0.0
    for i = 1:n
        for f in funs
            t += f(i)
        end
    end
    return t
end

@time test_elements(10^6) # 14.8k allocations
@time test_elements(10^6) # 6 allocations

@time test_loop(10^6) # 7 M allocations
@time test_loop(10^6) # 7 M allocations

```

---

<div class="post-metadata">

**Author:** ![JeffreySarnoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeffreysarnoff/32/1980_2.png) [@JeffreySarnoff](https://discourse.julialang.org/u/JeffreySarnoff)\
**Post date:** [May 22, 2019, 8:26pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/18 "2019-05-22T20:26:43Z")

</div>

Yes – `getindex` access works with “wrapped functions” in an array or a tuple.

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [May 22, 2019, 9:29pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/19 "2019-05-22T21:29:43Z")

</div>

As far as I understood, your arrays of functions are actually pretty small (~50) and you can probably hoist the dispatch: You could simply run a batch for each state. I.e. outer loop goes over states, inner loop over people.

If this is the case, then I recommend to just use a type for dispatch: Instead of storing a function, write `function foo(::State{:AZ}, params... )` for special cases and `function foo(::State{S}, params...) where S` for shared code. Plug in a function boundary, i.e. call a `@noinline batch_process(::State, data_batch)` and then your performance penalty should go away.

Or is there a strong reason that your outer loop must go over people / entities and your inner loop over states?

---

<div class="post-metadata">

**Author:** ![donboyd5](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/donboyd5/32/8562_2.png) [@donboyd5](https://discourse.julialang.org/u/donboyd5)\
**Post date:** [May 22, 2019, 9:54pm UTC](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471/20 "2019-05-22T21:54:29Z")

</div>

Thank you.

I’m sorry if I implied the inner loop in the example runs over multiple states. In this example, I intend to deal with just a single state. (Much, much later on, there might be multiple states.) The inner loop in my example is intended to go over a set of perhaps 30-50 functions that might apply for a single state.

If there is a way to address the question in that context, it would be really great. That is, trying to get rid of the performance penalty from looping through a set (tuple) of functions.

[Next page](https://discourse.julialang.org/t/array-of-functions-is-there-a-way-to-avoid-allocations-performance-penalty/24471.md?page=2)
