# Improving runtime

**URL:** <https://discourse.julialang.org/t/improving-runtime/101405>\
**Category:** New to Julia\
**Tags:** matlab, ode\
**Created:** [July 9, 2023, 4:58pm UTC](https://discourse.julialang.org/t/improving-runtime/101405 "2023-07-09T16:58:10Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Moslem\_Uddin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moslem_uddin/32/49483_2.png) [@Moslem\_Uddin](https://discourse.julialang.org/u/Moslem_Uddin)\
**Post date:** [July 9, 2023, 4:58pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/1 "2023-07-09T16:58:10Z")

</div>

I’m a new user of JULIA. The following code is a line by translation from MATLAB with the aim of comparing runtime. However, it seems like the performance doesn’t improve. I’ll be happy to have feedback on why it’s the case.

[https://github.com/muddin21/RegularizedStokeslet/blob/main/Regularized%20Stokeslet%20in%20JULIA.ipynb](https://github.com/muddin21/RegularizedStokeslet/blob/main/Regularized%20Stokeslet%20in%20JULIA.ipynb)

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [July 9, 2023, 5:09pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/2 "2023-07-09T17:09:48Z")

</div>

At first glance, it seems your code performs a lot of unneeded allocations. For Julia to be fast, you want to avoid allocating new memory when you can reuse it instead.  
See the [performance tips](https://docs.julialang.org/en/v1/manual/performance-tips/) for more details, especially the sections about

- [pre-allocating outputs](https://docs.julialang.org/en/v1/manual/performance-tips/#Pre-allocating-outputs)
- [fusing vectorized operations](https://docs.julialang.org/en/v1/manual/performance-tips/#More-dots:-Fuse-vectorized-operations)
- [using views for slices](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-views)

In many cases it will be as easy as replacing

```julia
for i in 1:n
    x = y - z
end

```

with

```julia
x = # init
for i in 1:n
    x .= y .- z
end

```

whenever you deal with arrays

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [July 9, 2023, 5:14pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/3 "2023-07-09T17:14:55Z")

</div>

You can diagnose the performance of your your code with the `@btime` macro from [BenchmarkTools.jl](https://github.com/JuliaCI/BenchmarkTools.jl). Ideally, you want the number of allocations not to scale with the number of loop iterations. In practice, that’s a lot to ask, so you should only focus on the most critical parts.  
How do I recognize these critical parts, you ask? By [profiling](https://docs.julialang.org/en/v1/manual/profile/) your code, eg. using the `@profview` macro from the [VSCode Julia extension](https://www.julia-vscode.org/docs/stable/userguide/profiler/).

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [July 9, 2023, 5:18pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/4 "2023-07-09T17:18:54Z")

</div>

Which function do you want to optimize?

---

<div class="post-metadata">

**Author:** ![fnin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fnin/32/23300_2.png) [@fnin](https://discourse.julialang.org/u/fnin)\
**Post date:** [July 11, 2023, 1:59pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/5 "2023-07-11T13:59:46Z")

</div>

As I understand it the main work is being done in the loop at the end of the notebook.

The loop is using non constant global variables. If you want Juila to be able to compile efficient code for the last part you should consider putting it in a function.

---

<div class="post-metadata">

**Author:** ![Moslem\_Uddin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moslem_uddin/32/49483_2.png) [@Moslem\_Uddin](https://discourse.julialang.org/u/Moslem_Uddin)\
**Post date:** [July 11, 2023, 2:52pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/6 "2023-07-11T14:52:05Z")

</div>

I would like to improve the overall performance of the code.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [July 11, 2023, 7:38pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/7 "2023-07-11T19:38:17Z")

</div>

Well, this is nothing we can help you with. We can help you with improving the performance of a function, if you point out which function should be faster.

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [July 11, 2023, 10:14pm UTC](https://discourse.julialang.org/t/improving-runtime/101405/8 "2023-07-11T22:14:25Z")

</div>

The advice I gave here should be a good starting point for you! Feel free to come back if there are things you don’t understand 🙂

> [@Improving runtime](https://discourse.julialang.org/t/improving-runtime/101405/2):
>
> At first glance, it seems your code performs a lot of unneeded allocations. For Julia to be fast, you want to avoid allocating new memory when you can reuse it instead. See the [performance tips](https://docs.julialang.org/en/v1/manual/performance-tips/) for more details, especially the sections about [pre-allocating outputs](https://docs.julialang.org/en/v1/manual/performance-tips/#Pre-allocating-outputs)[fusing vectorized operations](https://docs.julialang.org/en/v1/manual/performance-tips/#More-dots:-Fuse-vectorized-operations)[using views for slices](https://docs.julialang.org/en/v1/manual/performance-tips/#man-performance-views) In many cases it will be as easy as replacing for i in 1:n x = y - z end with x = # init for i in 1:n x .= y .- z end whenever you deal with arrays

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [July 12, 2023, 12:48am UTC](https://discourse.julialang.org/t/improving-runtime/101405/9 "2023-07-12T00:48:15Z")

</div>

One important thing I noticed is that you allocate an enormous number of 1-length vectors.

Now, this might make sense in Matlab (where everything is an array), but in Julia, `x=1` is very different from `x=[1]` in terms of memory allocation.

A small change from `X_s_distance2 = (X_s[1] .^ 2 .+ X_s[2] .^ 2)` to `X_s_distance2 = (X_s[1] ^ 2 + X_s[2] ^ 2)` alone is dropping the `@btime` results from 18 seconds to almost 16 seconds on my machine (yes, this `X_s_distance2` is one of your _one-element vectors_).

After doing this for a few more _one-element vectors_, I managed to reach 12 seconds - and I could continue doing this - but I think you got the point.

So, I don’t think your code is a line-by-line translation to Julia - especially because you are forcing this _everything is an array_ philosophy on Julia.

You might also be interested in reading [Noteworthy Differences from other Languages · The Julia Language](https://docs.julialang.org/en/v1/manual/noteworthy-differences/).

Equally important, pay attention to the advice related to performance that others have already pointed out. For example, in your `velocityRS` function, you allocate `u = zeros(2, length(s))` and `int_u = zeros(size(u))` each time when you call the function (and you end up calling it from a for loop). On my machine, preallocating `u` and `int_u` is followed by another almost 2 seconds execution time reduction. And there is still room to fix many things (e.g., an enormous number of allocations remain).

Have fun.

---

<div class="post-metadata">

**Author:** ![Moslem\_Uddin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/moslem_uddin/32/49483_2.png) [@Moslem\_Uddin](https://discourse.julialang.org/u/Moslem_Uddin)\
**Post date:** [July 14, 2023, 1:05am UTC](https://discourse.julialang.org/t/improving-runtime/101405/10 "2023-07-14T01:05:36Z")

</div>

Thank you for your comment. Even without pre-allocation, the runtime seems to be improved. Could you please show an example of the pre-allocation of one variable from the `velocityRS` function? I’m a little confused about this.

---

<div class="post-metadata">

**Author:** ![algunion](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/algunion/32/51630_2.png) [@algunion](https://discourse.julialang.org/u/algunion)\
**Post date:** [July 14, 2023, 10:55am UTC](https://discourse.julialang.org/t/improving-runtime/101405/11 "2023-07-14T10:55:21Z")

</div>

Please take a look [here](https://docs.julialang.org/en/v1/manual/performance-tips/#Pre-allocating-outputs) - this was already posted by @gdalle before.

This is one of the specific performance tips that you can apply to your `valocityRS` function: you allocate your output `u` each time you call the function inside your loop. Imagine the alternative where you preallocate `u` before your outer loop and then pass `u` to your `velocityRS` where you can fill/mutate the array as needed.

I am not saying this particular step will significantly improve your performance (because I think the big chunk of allocations are happening all over the place because of those unneeded 1-length vectors).

Another important thing - if you know that you will only need a very small container and its size is known at the compile time, you could use tuples instead of vectors (`(1,2)` vs. `[1,2]`), especially in the scenario of your nested loops - that will make a big difference on memory allocation (and garbage collection) - consequently will improve the execution.
