# Improving performance (Stochastic Path + Ipopt/Jump Optim on 32 cores/132RAM)

**URL:** <https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510>\
**Category:** Optimization (Mathematical)\
**Created:** [November 20, 2020, 5:35pm UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510 "2020-11-20T17:35:37Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Bruno](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bruno/32/18939_2.png) [@Bruno](https://discourse.julialang.org/u/Bruno)\
**Post date:** [November 20, 2020, 5:35pm UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/1 "2020-11-20T17:35:37Z")

</div>

Hello! I’ve created a [gist](https://gist.github.com/Brunobt80/0d19a78c012f460e8c1a2515d634a69c) with some Julia code. Basically I generate a path of a stochastic process (size 1000 array) and use it to analytically estimate one of the parameters (via nonlinear optim on a max likelihood surface) and repeat it (both generation/estimation) 500 times for each combination of parameters. I’ve implemented some naive strategies to run it in parallel (32 cpus) and improve performance. I noticed that the memory allocation is high (a couple of TiBs). Could someone please suggest any modifications and/or general advice that could improve the performance?

---

<div class="post-metadata">

**Author:** ![mikkoku](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikkoku/32/16274_2.png) [@mikkoku](https://discourse.julialang.org/u/mikkoku)\
**Post date:** [November 21, 2020, 6:40am UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/2 "2020-11-21T06:40:59Z")

</div>

I think parallelizing a single stochastic process is not useful, because the process is sequential. Is it possible to parallelize over the 500 replicates?

If you want to improve the performance, start with a single-threaded version and try to understand which function or which line of code is the slow one. Profiling helps in this, and I suggest using the profiler integrated with you IDE.

You can also try to observe the memory allocation of functions by using `@time` in suitable places.

It looks like the code has a lot of vector operations that will allocate something. In general it is better for performance to use loops instead of vector operations, unlike in matlab or R.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [November 21, 2020, 7:16am UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/3 "2020-11-21T07:16:46Z")

</div>

Your gist is a little too large for casual performance imrovement, I would try to boil it down to something you can post in a comment - it will get a better response.

---

<div class="post-metadata">

**Author:** ![Bruno](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bruno/32/18939_2.png) [@Bruno](https://discourse.julialang.org/u/Bruno)\
**Post date:** [November 21, 2020, 8:18pm UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/4 "2020-11-21T20:18:43Z")

</div>

Unfortunately the optimizer (Ipopt) always return an error (segmentation fault) whenever I try to parallelize it (even with just 2 threads). I’ll follow your advice and start with a single-threaded version. Thank you!

---

<div class="post-metadata">

**Author:** ![mikkoku](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikkoku/32/16274_2.png) [@mikkoku](https://discourse.julialang.org/u/mikkoku)\
**Post date:** [November 22, 2020, 8:10pm UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/5 "2020-11-22T20:10:42Z")

</div>

If you get segfaults with `@threads`, you might want to try [Distributed](https://docs.julialang.org/en/v1/stdlib/Distributed/) instead.

---

<div class="post-metadata">

**Author:** ![viralbshah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/viralbshah/32/54_2.png) [@viralbshah](https://discourse.julialang.org/u/viralbshah)\
**Post date:** [November 22, 2020, 9:27pm UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/6 "2020-11-22T21:27:30Z")

</div>

Could you file an issue at Ipopt.jl, with a reproducer for the segfault? It would be nice to figure out what the issue is.

-viral

---

<div class="post-metadata">

**Author:** ![cgeoga](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cgeoga/32/216186_2.png) [@cgeoga](https://discourse.julialang.org/u/cgeoga)\
**Post date:** [November 23, 2020, 3:37am UTC](https://discourse.julialang.org/t/improving-performance-stochastic-path-ipopt-jump-optim-on-32-cores-132ram/50510/7 "2020-11-23T03:37:51Z")

</div>

I’m not having much luck running this code. For example, if I comment out everything in the function `run` after the line `df = hcat(...)` and then call `run()`, I get an error that the two things being `hcat`-ed are not compatible in shape. It’s possible that some part of your code is not working in a more basic way but for some reason you’re just not getting the most helpful error information, but whatever’s going on is plausibly not really because of threading or Ipopt.

As has already been suggested, though, the gist is a little hard to read with a specific eye to parallelism issues. I only see two `Threads.@threads` annotations, and they both seem to me to be on loops that should be over 1000 elements. What would be more helpful for us to look at, and would probably result in a more efficient program execution, would be to write a compartmentalized function that does the simulation and estimation, and then just toss that into a `ThreadsX.map` or `pmap` or something and see in that case if the speedup isn’t as good as you expect.

EDIT: oh, geez, I didn’t see that you already opened an example with a different MWE on the Ipopt page. Whoops.
