# \[ANN\] ArrayAllocators.jl v0.3 composes with OffsetArrays.jl v1.12.1+ for faster zeros with offset indexing

**URL:** <https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161>\
**Category:** Package Announcements\
**Tags:** announcement\
**Created:** [June 22, 2022, 5:04am UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161 "2022-06-22T05:04:29Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [June 22, 2022, 5:04am UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161/1 "2022-06-22T05:04:29Z")

</div>

[ArrrayAllocators.jl v0.3.0+](https://github.com/mkitti/ArrayAllocators.jl/releases/tag/v0.3.0) and [OffsetArrays.jl v1.12.1+](https://github.com/JuliaArrays/OffsetArrays.jl/releases/tag/v1.12.1) now compose. This means that you can now construct an `OffsetArray` of `0`s using the following equivalent lines of code.

```julia
julia> using OffsetArrays, ArrayAllocators

julia> OA = OffsetArray{Int}(undef, -512:511); fill!(OA, 0);

julia> OA_calloc = OffsetArray{Int}(calloc, -512:511);

julia> OA[-512] == OA_calloc[-512]
true

julia> isequal(OA_calloc, OA)
true

```

In some cases, this can yield a reduction in the initialization time of the `OffsetArray` as illustrated below.

```julia
julia> using BenchmarkTools

julia> @btime begin
           OA2 = OffsetArray{Int}(undef, -512:511, -512:511)
           fill!(OA2, 0)
       end;
  1.483 ms (2 allocations: 8.00 MiB)

julia> @btime begin
           OA2 = OffsetArray{Int}(calloc, -512:511, -512:511)
       end;
  1.089 ms (8 allocations: 8.00 MiB)

julia> @btime begin
           fill!(OA2, 1)
       end setup = (OA2 = OffsetArray{Int}(undef, -512:511, -512:511));
  1.521 ms (0 allocations: 0 bytes)

julia> @btime begin
           fill!(OA2_calloc, 1)
       end setup = (OA2_calloc = OffsetArray{Int}(calloc, -512:511, -512:511));
  1.312 ms (0 allocations: 0 bytes)

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/c/a/ca0153c511ff90a3a214277bb7df0431c2767f61.png)

I encourage you to vigorously benchmark your application when using this as the potential optimization may be operating system and hardware dependent.

Also with the subpackage NumaAllocators v0.2.0, you may now directly allocate an `OffsetArray` on a specific non-uniform memory architecture (NUMA) node.

```julia
julia> using OffsetArrays, NumaAllocators

julia> OA_numa_0 = OffsetArray{Int}(numa(0), -512:512, -9:9);

```

The composition was enabled by implementing `Base.unsafe_wrap` for `OffsetArray`. Neither package depends directly upon the other, so ensure that both packages are at the required version or later.

```julia
julia> ptr = Libc.calloc(1024, 8)
Ptr{Nothing} @0x0000000004e3ae10

julia> OA3 = unsafe_wrap(OffsetArray, Ptr{Int}(ptr), 1024);

```

> <https://github.com/JuliaArrays/OffsetArrays.jl/pull/290>
>
> Fix #275

If you missed the original ArrayAllocators.jl announcement, see below.

> [@\[ANN\] ArrayAllocators.jl: Integrating calloc and aligned memory into Array construction](https://discourse.julialang.org/t/ann-arrayallocators-jl-integrating-calloc-and-aligned-memory-into-array-construction/80356):
>
> ArrayAllocators.jl I am happy to announce [ArrayAllocators.jl](https://github.com/mkitti/ArrayAllocators.jl), a registered package that provides new mechanisms of array allocation. ArrayAllocators.jl provides new values that can take the place of undef when constructing arrays via the Array constructor: Array{T}(allocator, n, m, ...). Quick Start Example For example, you can now do the following. using ArrayAllocators malloced\_array = Array{Int}(malloc, 16, 32, 8) calloced\_array = Array{Int}(calloc, 1024) aligned\_array = Array{Int}(MemAlign…

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [June 22, 2022, 6:17am UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161/2 "2022-06-22T06:17:37Z")

</div>

Another feature of ArrayAllocators.jl v0.3 is an alternate implementation of `zeros` via `calloc`. Note that this `ArrayAllocators.zeros` is not exported. Here `ArrayAllocators.zeros(T, ...)` is essentially just `Array{T}(calloc, ...)`.

```julia
julia> using ArrayAllocators

julia> @time AAZ = ArrayAllocators.zeros(Int, 1024, 1024, 1024);
  0.029364 seconds (61.31 k allocations: 8.004 GiB, 99.89% compilation time)

julia> @time AAZ = ArrayAllocators.zeros(Int, 1024, 1024, 1024);
  0.000037 seconds (4 allocations: 8.000 GiB)

julia> @time BZ = Base.zeros(Int, 1024, 1024, 1024);
  4.448959 seconds (2 allocations: 8.000 GiB, 0.52% gc time)

julia> @time BZ = Base.zeros(Int, 1024, 1024, 1024);
  4.665584 seconds (2 allocations: 8.000 GiB, 2.48% gc time)

```

Note that on some operating systems this may defer the actual allocation of memory until writing.

```julia
julia> @time fill!(AAZ, 1)
  4.849603 seconds (8.17 k allocations: 491.429 KiB, 0.64% compilation time)

julia> @time fill!(AAZ, 2);
  1.634943 seconds

julia> @time fill!(BZ, 1);
  1.710879 seconds

julia> @time fill!(BZ, 2);
  1.717882 seconds

```

If you wanted to switch your code to use `ArrayAllocators.zeros` instead of `Base.zeros`, you can import it before your first reference to `zeros`.

```julia
julia> using ArrayAllocators: zeros

julia> @time zeros(Int, 1024, 1024, 1024);
  0.000028 seconds (4 allocations: 8.000 GiB)

julia> @time Base.zeros(Int, 1024, 1024, 1024);
  4.442414 seconds (2 allocations: 8.000 GiB, 0.10% gc time)

```

While this may look like a free lunch, note @Sukera’s caveats in the original post:

> [@Faster zeros with calloc](https://discourse.julialang.org/t/faster-zeros-with-calloc/69860):
>
> Abstract TL; DR zeros\_via\_calloc is a potentially faster version of zeros that is comparable to Array{T}(undef, ...) and numpy.zeros. function zeros\_via\_calloc(::Type{T}, dims::Integer...) where T ptr = Ptr{T}(Libc.calloc(prod(dims), sizeof(T))) return unsafe\_wrap(Array{T}, ptr, dims; own=true) end # Windows benchmark julia\> @btime zeros\_via\_calloc(Float64, 1024, 1024); 12.400 μs (2 allocations: 8.00 MiB) julia\> @btime zeros(Float64, 1024, 1024); 1.652 ms (2 allocations: 8.00 MiB)…

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 29, 2022, 7:19am UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161/3 "2022-06-29T07:19:53Z")

</div>

Your package was a cool read! And as I somewhat picked up from the reading, calloc makes the bits 0 but not the elements `zero`, so I’ll watch out for that:

```julia
julia> using ArrayAllocators

julia> struct Zerone{T<:Real}
         parent::T
       end

julia> Base.zero(::Type{Zerone{T}}) where T = Zerone{T}(oneunit(T))

julia> Z = zeros(Zerone{Float64}, 2, 2)
2×2 Matrix{Zerone{Float64}}:
 Zerone{Float64}(1.0) Zerone{Float64}(1.0)
 Zerone{Float64}(1.0) Zerone{Float64}(1.0)

julia> U = Array{Zerone{Float64}}(calloc, 2, 2)
2×2 Matrix{Zerone{Float64}}:
 Zerone{Float64}(0.0) Zerone{Float64}(0.0)
 Zerone{Float64}(0.0) Zerone{Float64}(0.0)

julia> Za = ArrayAllocators.zeros(Zerone{Float64}, 2, 2)
2×2 Matrix{Zerone{Float64}}:
 Zerone{Float64}(0.0) Zerone{Float64}(0.0)
 Zerone{Float64}(0.0) Zerone{Float64}(0.0)

```

I do have a question, though, if the array allocation is only deferred until next setindexing, am I right in thinking that alone would not improve performance? That the only scenario where I can save allocation time is if my OS has preallocated pages of 0s?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [June 29, 2022, 7:32am UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161/4 "2022-06-29T07:32:21Z")

</div>

> [@Benny](#):
>
> Your package was a cool read! And as I somewhat picked up from the reading, calloc makes the bits 0 but not the elements `zero`, so I’ll watch out for that:

Indeed - if “blob of zeros” is not a valid representation for your type, just using `calloc` will not help you.

> [@Benny](#):
>
> I do have a question, though, if the array allocation is only deferred until next setindexing, am I right in thinking that alone would not improve performance? That the only scenario where I can save allocation time is if my OS has preallocated pages of 0s?

Sort of. In general it depends on how your OS is handling calls to `calloc` in the libc you’re using. Linux at least has a dedicated zeroed page, to allow reads from `calloc`ed memory to be fast, but this of course doesn’t help with writes - the first write still has to allocate a page, which takes time. So in effect, the time an allocation takes is moved from the `calloc` call to the first real use of the memory.

There’s a bunch of discussion in this related thread:

> [@Faster zeros with calloc](https://discourse.julialang.org/t/faster-zeros-with-calloc/69860):
>
> Abstract TL; DR zeros\_via\_calloc is a potentially faster version of zeros that is comparable to Array{T}(undef, ...) and numpy.zeros. function zeros\_via\_calloc(::Type{T}, dims::Integer...) where T ptr = Ptr{T}(Libc.calloc(prod(dims), sizeof(T))) return unsafe\_wrap(Array{T}, ptr, dims; own=true) end # Windows benchmark julia\> @btime zeros\_via\_calloc(Float64, 1024, 1024); 12.400 μs (2 allocations: 8.00 MiB) julia\> @btime zeros(Float64, 1024, 1024); 1.652 ms (2 allocations: 8.00 MiB)…

but I personally don’t think `calloc` is a good idea, since it can mask performance problems due to the time being spent no longer being directly linked to the original allocation.

---

<div class="post-metadata">

**Author:** ![mkitti](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkitti/32/12459_2.png) [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Post date:** [June 30, 2022, 8:12pm UTC](https://discourse.julialang.org/t/ann-arrayallocators-jl-v0-3-composes-with-offsetarrays-jl-v1-12-1-for-faster-zeros-with-offset-indexing/83161/5 "2022-06-30T20:12:41Z")

</div>

> [@Benny](#):
>
> I do have a question, though, if the array allocation is only deferred until next setindexing, am I right in thinking that alone would not improve performance? That the only scenario where I can save allocation time is if my OS has preallocated pages of 0s?

My recommedation is to throughly benchmark your application. ArrayAllocators.jl makes it easier to switch allocation methods by providing a mechanism through the standard syntax.

Yes, the benchmarks will be confounded as some write operations will include allocation time.

The time savings will depend how you write. Allocation is not necessarily an all or none situation. Your operating system allocates “pages” of memory at a time. If you are not going to write to the entire array, there are potential savings. If you are going to write to all elements of the array, then this will not be very useful over `undef`.
