# Discussion on the future for performant programming in Julia

**URL:** <https://discourse.julialang.org/t/discussion-on-the-future-for-performant-programming-in-julia/99509>\
**Category:** Internals & Design\
**Tags:** question\
**Created:** [May 28, 2023, 10:29am UTC](https://discourse.julialang.org/t/discussion-on-the-future-for-performant-programming-in-julia/99509 "2023-05-28T10:29:20Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [May 28, 2023, 1:48pm UTC](https://discourse.julialang.org/t/discussion-on-the-future-for-performant-programming-in-julia/99509/4 "2023-05-28T13:48:37Z")

</div>

> [@jt1](#):
>
> Once differentiation is required, it requires [400 lines of code](https://github.com/SciML/SciMLSensitivity.jl/blob/834a4c70c45c6525d66d6fa7ee7c29837435e0f9/src/forward_sensitivity.jl) just to get forward-mode AD with memory management to work efficiently.

Nope. Correction, that’s for forward sensitivity analysis, which is equivalent to forward-mode AD (with a few differences in computation order). Forward-mode AD of the solver works without any code from SciMLSensitivity, which is why it’s performed even if that package isn’t loaded.

> [@jt1](#):
>
> In the prime example of high performance computing, [DifferentialEquations.jl](https://sciml.ai/), all required memory is cached beforehand like [here](https://github.com/SciML/OrdinaryDiffEq.jl/blob/master/src/caches/rosenbrock_caches.jl)

It’s very common in HPC to have strong memory bounds and have to know the exact memory requirement. This way we can enforce strong real-time guarantees on the amount of memory that is required by the methods. You can see this in action in the low-storage RK methods which actually test this property to ensure we have exactly the theoretical number of arrays allocated.

> <https://github.com/SciML/OrdinaryDiffEq.jl/blob/b7bb73f46669e6f35c33baf2b4a1a1dc9892a091/test/algconvergence/ode_low_storage_rk_tests.jl#L109>

Of course, DifferentialEquations.jl is used in a lot of real-time applications where guaranteeing all memory is allocated up front to have an essentially static stepper is a requirement. So even with improvements, DiffEq has some applications which PyTorch’s model cannot support which we want to continue to support, and even support better (going all the way to static compilation and edge computing). Therefore it’s not going to go away as this code is more similar to C++ or Rust in terms of the application’s goals.

That said.

> [@jt1](#):
>
> This approach takes too much time, is too hard for most developers and doesn’t scale to more heavy array calculations like in Machine learning. I would love to know if Julia will tackle this and how and when.
> 
> The difference with `Pytorch` which doesn’t have these problems is that Julia arrays don’t know their shape at compile-time and hence automatic memory management of arrays isn’t possible. I think having no separation between `lists` and `arrays` in Julia is a mistake, they should share a supertype though. **This would be solved by adding arrays with known shape to Base and writing automatic array memory management for them** and having AD work well with them. I Think `MArray`s from `StaticArrays.jl` is a good starting point. But with a name like `ShapedArrays` for example.
> 
> **With this, it would be possible to statically compile most type stable code without shipping a GC** , since most type stable functions only allocate due to Arrays. Moreover, compilation to MLIR, GPU and others should be doable.

This is all being discussed and there’s a small PRs in motion. v1.9 and v1.10 had a large focus on TTFX and usability issues, with the end of v1.10 getting some things like improved stacktraces a major push (for example, improved stacktraces for automatic differentiation [Frame registration for foreign JITs · Issue #49056 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/49056) and generally truncating stacktraces [https://twitter.com/ChrisRackauckas/status/1661014235466563591](https://twitter.com/ChrisRackauckas/status/1661014235466563591)). With those kinds of usability improvements and loading improvements, v1.10 is set to branch off its release rather soon.

With those changes in hand, we have been having many discussions about improving the way that high performance and statically compilable code can be written. This includes steps like:

- Improving the performance of the allocator ([memory reuse between iterations · Issue #49881 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/49881) and [Allocating empty array is 4x slower than Java · Issue #45160 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/45160) are example MWEs)
- Making the GC format faster to free
- Improving escape analysis and shape inference so that the compiler can optimize with known lifetimes
- Changing the implementation of arrays to be based on a generic memory buffer, which other types can then build off of.

The MWEs came from a recent priorities meeting so while it’s hard to put a timeline on things, I can at least say that the improvement of memory management on performance has now hit the top of the stack in terms of some JuliaHub priorities and I would expect a major focus in this area, along with static compilation, in the near future.

We will (hopefully) be writing a Julep in the very near future that details some motivations and some planned routes to solve some issues along these lines, so I’ll just cut here for now.

---

_[View the full topic](https://discourse.julialang.org/t/discussion-on-the-future-for-performant-programming-in-julia/99509)._
