# Improved allocation design, with 4-byte pointers, and sometimes 5-byte in effect

**URL:** <https://discourse.julialang.org/t/improved-allocation-design-with-4-byte-pointers-and-sometimes-5-byte-in-effect/123229>\
**Category:** Internals & Design\
**Created:** [November 28, 2024, 9:29pm UTC](https://discourse.julialang.org/t/improved-allocation-design-with-4-byte-pointers-and-sometimes-5-byte-in-effect/123229 "2024-11-28T21:29:10Z")\
**Posts on this page:** 1\
**Showing post:** 9

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [December 9, 2024, 5:49pm UTC](https://discourse.julialang.org/t/improved-allocation-design-with-4-byte-pointers-and-sometimes-5-byte-in-effect/123229/9 "2024-12-09T17:49:15Z")

</div>

@Oscar_Smith, @GeorgeGkountouras Immix GC (./julia) is confidently 33% faster than Julia 1.11 (with its default GC), for _three_ threads, which is the optimal number for at least than program with (at least) Immix, on this worst-case outlier from Debian Benchmark Game, though only 11% faster than 1.10. Julia 1.11 (default GC) is 17% slower than that Julia 1.10 (default GC), i.e. Immix (1.12.0-DEV.1745) is 33% faster than 1.11.

Since the benchmark game uses four threads (all the cores), and I believe insists on the same config, `-t 4`, for all programs, is there a way to opt into _fewer_ at runtime? I know you can’t yet, ask for more at runtime (except by a hack, by calling C, and it adds…):

[https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/binarytrees-julia-3.html](https://benchmarksgame-team.pages.debian.net/benchmarksgame/program/binarytrees-julia-3.html)

```julia
~/MMTk/julia$ killall -SIGSTOP firefox-bin

 hyperfine './julia -t 3 binarytrees.julia-3.julia 21'
Benchmark 1: ./julia -t 3 binarytrees.julia-3.julia 21
  Time (mean ± σ): 7.488 s ± 0.203 s [User: 18.221 s, System: 0.438 s]
  Range (min … max): 7.215 s … 7.781 s 10 runs

$ hyperfine 'julia -t 3 binarytrees.julia-3.julia 21'
Benchmark 1: julia -t 3 binarytrees.julia-3.julia 21
  Time (mean ± σ): 8.877 s ± 0.501 s [User: 13.945 s, System: 3.693 s]
  Range (min … max): 8.144 s … 9.498 s 10 runs

$ hyperfine 'julia +1.11 -t 3 binarytrees.julia-3.julia 21'
Benchmark 1: julia +1.11 -t 3 binarytrees.julia-3.julia 21
  Time (mean ± σ): 10.254 s ± 0.306 s [User: 16.358 s, System: 3.016 s]
  Range (min … max): 9.609 s … 10.585 s 10 runs

I was accidentally benchmarking a non-threaded program (for Distributed) before, and then Immix was is confidently slower than Julia 1.10's GC (defaults 1.11's GC is also slower):

~/MMTk/julia$ hyperfine './julia -t 16 ../../binarytrees.julia-4.julia 21'
Benchmark 1: ./julia -t 16 ../../binarytrees.julia-4.julia 21
  Time (mean ± σ): 11.998 s ± 0.133 s [User: 28.054 s, System: 0.724 s]
  Range (min … max): 11.777 s … 12.237 s 10 runs
 
$ hyperfine './julia -t 4 ../../binarytrees.julia-4.julia 21'
Benchmark 1: ./julia -t 4 ../../binarytrees.julia-4.julia 21
  Time (mean ± σ): 11.586 s ± 0.120 s [User: 16.106 s, System: 0.355 s]
  Range (min … max): 11.454 s … 11.804 s 10 runs
 
$ hyperfine './julia ../../binarytrees.julia-4.julia 21'
Benchmark 1: ./julia ../../binarytrees.julia-4.julia 21
 ⠼ Current estimate: 12.070 s ██████████████████████████████████████████████████████████████████████████████████████████████████████████████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ETA 00:00:49
^C
$ hyperfine 'julia ../../binarytrees.julia-4.julia 21'
Benchmark 1: julia ../../binarytrees.julia-4.julia 21
 ⠼ Current estimate: 10.995 s ██████████████████████████████████████████████████████████████████████████████████████████████████████████████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ETA 00:00:44
^C

Startup is a bit slower with Immix:
$ hyperfine './julia -e ""'
Benchmark 1: ./julia -e ""
  Time (mean ± σ): 232.5 ms ± 17.5 ms [User: 309.9 ms, System: 82.7 ms]
  Range (min … max): 203.5 ms … 248.3 ms 12 runs
 

$ hyperfine 'julia +1.11 -e ""'
Benchmark 1: julia +1.11 -e ""
  Time (mean ± σ): 192.3 ms ± 14.5 ms [User: 281.3 ms, System: 65.8 ms]
  Range (min … max): 166.9 ms … 210.2 ms 14 runs

$ hyperfine 'julia +1.10 -e ""'
Benchmark 1: julia +1.10 -e ""
  Time (mean ± σ): 225.5 ms ± 18.5 ms [User: 207.2 ms, System: 141.6 ms]
  Range (min … max): 198.3 ms … 247.1 ms 12 runs

$ ./julia
  | | |_| | | | (_| | | Version 1.12.0-DEV.1745 (2024-12-09)
 _/ |\ __'_|_|_|\__'_| | upstream-ready/immix/4aeea613eb (fork: 27 commits, 8 days)

```

Side-note, while compiling I saw:

> Precompiling packages ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╸━━━ 98/106  
> ◓ Pkg -g2 -O3  
> ◒ REPL -g2 -O3  
> ◓ Pkg -g2 --check-bounds=yes -O3  
> ◒ REPL -g2 --check-bounds=yes -O3

I believe we have duplicated pkgimages for that, and we could have just one for the most conservative, with `--check-bounds=yes` at least for those that are not speed critical.

EDIT2: Compiling Immix as we speak, I unstuck the segfault problem, by changing:

$ (cd julia && git checkout dev && echo ‘MMTK\_PLAN=Immix’ \> Make.user)

to:

$ (cd julia && git checkout upstream-ready/immix && echo ‘MMTK\_PLAN=Immix’ \> Make.user) # upstream-ready/immix

[Some more info on how to compile in history for this post, since my experiments in getting to compile, and fixing the segfault in building Julia, are a distraction here.]

EDIT: ~~if someone wants to help with building julia~~ , it segfaults… building MMtk in previous step worked, but wasn’t used below, since you need to build MMTk AND then julia with it also from source:

I was hoping even faster with the new ~~upcoming GC~~ PR (it’s disabled here, since I wasn’t building from source), but it is 4.8% faster, though only compared to 1.11 (for min):

```julia
$ juliaup add pr56288

$ hyperfine 'julia +1.11 -t 4 binarytrees.julia-4.julia 21'
Benchmark 1: julia +1.11 -t 4 binarytrees.julia-4.julia 21
  Time (mean ± σ): 12.591 s ± 0.304 s [User: 14.098 s, System: 0.530 s]
  Range (min … max): 12.248 s … 13.252 s 10 runs
 
$ hyperfine 'julia +pr56288 -t 4 binarytrees.julia-4.julia 21'
Benchmark 1: julia +pr56288 -t 4 binarytrees.julia-4.julia 21
  Time (mean ± σ): 12.059 s ± 0.246 s [User: 16.355 s, System: 0.559 s]
  Range (min … max): 11.654 s … 12.539 s 10 runs

```

both are even regressions from 1.10, even when it’s single threaded:

```julia
$ hyperfine 'julia +1.10 binarytrees.julia-4.julia 21'
Benchmark 1: julia +1.10 binarytrees.julia-4.julia 21
  Time (mean ± σ): 11.352 s ± 0.244 s [User: 11.430 s, System: 0.765 s]
  Range (min … max): 11.068 s … 11.925 s 10 runs
 
$ hyperfine 'julia +1.10 -t 4 binarytrees.julia-4.julia 21'
Benchmark 1: julia +1.10 -t 4 binarytrees.julia-4.julia 21
 ⠏ Current estimate: 10.957 s █████████████████████████████████████████████████████████░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ETA 00:01:17
^C

```

I’m assuming Immix is enabled by default after compiling from source, while I did have to change this like (building Julia as is with it segfaulted):

```julia
(cd julia && git checkout dev && echo 'MMTK_PLAN=Immix' > Make.user) # or MMTK_PLAN=StickyImmix to use Sticky Immix

```

to:

```julia
(cd julia && git checkout upstream-ready/immix && echo 'MMTK_PLAN=Immix' > Make.user) # upstream-ready/immix

```

I’ve yet to try out StickyImmix, do you know the difference?

I’m looking into the new MMTk, ~~and if I need to tune it somehow~~ :

> <https://github.com/JuliaLang/julia/pull/56288>
>
> This PR adds the possibility of building/running Julia using MMTk, running non-m…oving immix.
> The binding code (Rust) associated with it is in https://github.com/mmtk/mmtk-julia/tree/upstream-ready/immix. Instructions on how to build/run Julia with MMTk are described in the \[\`README\`\](https://github.com/mmtk/mmtk-julia/tree/upstream-ready/immix?tab=readme-ov-file#an-mmtk-binding-for-the-julia-programming-language) file inside the binding repo.

I’m built with the instructions from there (modified as explained above):

> **[GitHub - mmtk/mmtk-julia at upstream-ready/immix](https://github.com/mmtk/mmtk-julia/tree/upstream-ready/immix)**
>
> upstream-ready/immix

first:

```julia
$ sudo apt install cargo

```

> For example, MMTk provides BumpPointer, which simply includes a cursor and a limit.In the following example, we embed one BumpPointer struct in the TLS.

Does Julia itself need to opt into something like: [post\_alloc in mmtk::memory\_manager - Rust](https://docs.mmtk.io/api/mmtk/memory_manager/fn.post_alloc.html)

I see in the code:

> INLINE\_FASTPATH\_ALLOCATION

Something else to look into and might be related to what I’m proposing:

> [@Should Dict avoid GC allocations for isbits key/value types?](https://discourse.julialang.org/t/should-dict-avoid-gc-allocations-for-isbits-key-value-types/123613):
>
> Resizing arrays, when elements have primitive types or more generally isbits types, does not trigger any GC, as this is handled by the C runtime. Dictionaries are also allocation-free during most operations, except during rehashing, which happens if you insert more elements than the current capacity. rehash is called to allocate larger arrays inside the Dict object before data from the original arrays are copied over. The original arrays are then discarded and eventually garbage-collected. It s…

---

_[View the full topic](https://discourse.julialang.org/t/improved-allocation-design-with-4-byte-pointers-and-sometimes-5-byte-in-effect/123229)._
