# Blog: Using Julia on the HPC

**URL:** <https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039>\
**Category:** Teaching & Outreach\
**Tags:** blog-post\
**Created:** [September 30, 2022, 1:17pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039 "2022-09-30T13:17:43Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [September 30, 2022, 1:17pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/1 "2022-09-30T13:17:43Z")

</div>

Hi all!

Recently, I gave a talk at a HPC conference on using Julia to take advantage of the low barrier to entry for high-performance parallel code at every level (threading, multiprocessing, GPU etc). I was asked to write a follow up blog post on the same topic. While I do not talk about everything, it may be of interest - you can view the [blog post here](https://blogs.nottingham.ac.uk/digitalresearch/2022/09/30/using-julia-on-the-hpc/).

There is also an accompanying GitHub repo with all the code [here](https://github.com/JamieMair/julia-for-research-with-hpc).

p.s. In the blog, I do a performance comparison with C++, but I am no C++ programmer, and it would be useful if someone could comment on whether I am being fair in the comparison.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [September 30, 2022, 5:32pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/3 "2022-09-30T17:32:26Z")

</div>

I really wish ClusterManagers.jl is more robust and self-contained and has multiple way of connecting worker/main node under different network conditions

---

<div class="post-metadata">

**Author:** ![MaciekBielski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maciekbielski/32/14225_2.png) [@MaciekBielski](https://discourse.julialang.org/u/MaciekBielski)\
**Post date:** [September 30, 2022, 8:36pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/4 "2022-09-30T20:36:41Z")

</div>

I am a complete rookie in Julia but know a bit more about C++. I have skimmed through the blog post and related GitHub project, I found them quite interesting, thanks for your effort!  
My only question is related to optimisation levels, at least for CPU execution. Did you use defaults or maximum levels in both cases? IMHO, the latter would be a really fair comparison of perf capabilites between C++ vs. Julia.

---

<div class="post-metadata">

**Author:** ![gbellomia](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbellomia/32/19443_2.png) [@gbellomia](https://discourse.julialang.org/u/gbellomia)\
**Post date:** [September 30, 2022, 8:59pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/5 "2022-09-30T20:59:53Z")

</div>

A curiosity:

 ![julia_code_5-1024x547~3](https://global.discourse-cdn.com/julialang/original/3X/9/e/9ec7c9aff19d62a3344c2ccb07b27c21221ffd2d.png)

Why do you need to fill with zeros inside the function, if you are initializing to all zeros before the call? (and later you do the same by initializing the CUDA array and the Darray).

Am I missing something?

Anyway very nice article, I enjoyed the read!

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [September 30, 2022, 9:00pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/6 "2022-09-30T21:00:33Z")

</div>

Thanks for the interest.

For the C++ compilation I just used `g++ -O3 main.cpp`. Is there anything that I missed?

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [September 30, 2022, 9:15pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/7 "2022-09-30T21:15:40Z")

</div>

Ah yes! That’s a good point, you haven’t missed anything. In my original talk, I moved all allocations outside the benchmarked functions, but I chose to include the allocation line in the blog post, as I thought it was easier to understand. So having the input set to zero was a relic from the first iteration. I suppose I wanted `random_walk!` to work with “unzeroed” memory.

Hopefully, resetting to zero doesn’t bias the results too much (especially when `T=100` as in the benchmarks), but a keen observation nonetheless!

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [September 30, 2022, 9:18pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/8 "2022-09-30T21:18:56Z")

</div>

Yes, it seems like it wasn’t straightforward for some of my colleagues to use it.

I have a script which handles the cluster manager stuff using the environment variables set by SLURM which takes a main script and an “include” file to run on each worker before the main script, but it would be nice if I didn’t need this.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 1, 2022, 3:53am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/9 "2022-10-01T03:53:48Z")

</div>

You are simply comparing the speed of RNG between two languages. Based on my personal experience, Julia threading is never as fast as OpenMP.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 1, 2022, 3:54am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/10 "2022-10-01T03:54:48Z")

</div>

> [@Yifan\_Liu](#):
>
> Based on my personal experience, Julia threading is never as fast as OpenMP.

fast in what sense? you may want [GitHub - JuliaSIMD/Polyester.jl: The cheapest threads you can find!](https://github.com/JuliaSIMD/Polyester.jl) if you’re looking for low-overhead kind of “fast”?

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 1, 2022, 3:57am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/11 "2022-10-01T03:57:53Z")

</div>

There is no benchmark against OpenMP there. Also, if it is so good, then why not replace Threads.@threads with it?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [October 1, 2022, 3:58am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/12 "2022-10-01T03:58:42Z")

</div>

> [@Yifan\_Liu](#):
>
> There is no benchmark against OpenMP there.

what do YOU mean by “fast” then, what benchmark are you even looking for.

> [@Yifan\_Liu](#):
>
> if it is so good, then why not replace Threads.@threads with it?

if OpenMP is so good why are you using Julia /jk

* * *

`Base` stuff (Threads stdlib) needs to strike some kind of balance, Polyester is good for low over-head stuff when you have particularly small tasks

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 1, 2022, 4:03am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/13 "2022-10-01T04:03:30Z")

</div>

> [@jling](#):
>
> if OpenMP is so good why are you using Julia /jk

I no longer use Julia. I have migrated back to R/Python and C++. OpenMP is really a masterpiece. Not sure if Julia threading has caught up.

---

<div class="post-metadata">

**Author:** ![MaciekBielski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maciekbielski/32/14225_2.png) [@MaciekBielski](https://discourse.julialang.org/u/MaciekBielski)\
**Post date:** [October 1, 2022, 7:07am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/14 "2022-10-01T07:07:35Z")

</div>

No, at least for C++ part that really sets the baseline.

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [October 1, 2022, 8:24am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/15 "2022-10-01T08:24:16Z")

</div>

In my experience, my workloads are usually big enough that the overhead of Julia multithreading becomes negligible. If it is a problem, there are alternatives to the base implementation which are faster and more suitable (as mentioned by @jling).

I don’t know when you last used Julia, but the Threads library has been improving on every Julia release, so maybe it’s worth a try to compare overheads with OpenMP, I would be interested to see that benchmark.

I introduce threading here, as threading is not even available in languages like Python and MATLAB. I only include the C++ comparison to “hook” some people in, who may be only be interested only if Julia is fast like C, not slow like Python.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [October 1, 2022, 8:29am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/16 "2022-10-01T08:29:05Z")

</div>

> [@jmair](#):
>
> I only include the C++ comparison to “hook” some people in, who may be only be interested only if Julia is fast like C, not slow like Python.

But this comparison is incorrect. Julia uses the Xoshiro256++ algorithm for its RNG.

---

<div class="post-metadata">

**Author:** ![jmair](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmair/32/35117_2.png) [@jmair](https://discourse.julialang.org/u/jmair)\
**Post date:** [October 1, 2022, 9:31am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/17 "2022-10-01T09:31:39Z")

</div>

A valid point, however, I am simply comparing the defaults, as I said, what would be likely from a new PhD student just trying to get some simulations working. My main aim, is not to show that it is “faster”, just that in Julia, it is easy to reach very high levels of performance for very little effort.

I think this is of interest to programmers and researchers that have only used Python or MATLAB, and would find learning Julia more appealing than C/C++ etc, knowing that they aren’t forced into using those languages if they want parallelism or high performance.

---

<div class="post-metadata">

**Author:** ![greatpet](https://avatars.discourse-cdn.com/v4/letter/g/e495f1/32.png) [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Post date:** [October 1, 2022, 1:47pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/18 "2022-10-01T13:47:10Z")

</div>

> [@Yifan\_Liu](#):
>
> OpenMP is really a masterpiece. Not sure if Julia threading has caught up.

If I have threads which also create threads, is OpenMP as easy as Julia if you want the threads to be scheduled so that you never run more threads than the number of CPU cores? (Parallel computing noob here. Would love to read more about comparisons of the various options like OpenMP, IntelTBB, Polyester.jl etc.)

---

<div class="post-metadata">

**Author:** ![gbellomia](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gbellomia/32/19443_2.png) [@gbellomia](https://discourse.julialang.org/u/gbellomia)\
**Post date:** [October 4, 2022, 11:52pm UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/19 "2022-10-04T23:52:52Z")

</div>

> [@jmair](#):
>
> threading is not even available in languages like Python and MATLAB

I’m sorry, but this is plain false. Matlab automatically exploits threading on many many builtin functionality (ffts, matrix products, generic array operations…): you may want to take a look on their [dedicated page](https://it.mathworks.com/discovery/matlab-multicore.html).

You could argue that a serious user would want fine control over what’s actually happening (although I’d respond that big part of matlab’s success has always been about its _very_ good default handling of many things, from algorithm choice to these kind of things, to wisdom for fftw stuff and so on… so yes, but you’d better be a true expert to actually beat their careful default setup). In that case you just need to invoke one of the many constructs provided by the parallel computing toolbox, where you can go from a naive `parfor` loop to a fine control over the number of workers and how distributed memory is managed. Similarly you can easily initialize an array on the GPU and let Matlab dispatch therein all your array operations.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [October 5, 2022, 12:37am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/20 "2022-10-05T00:37:57Z")

</div>

> [@gbellomia](#):
>
> I’m sorry, but this is plain false. Matlab automatically exploits threading on many many builtin functionality (ffts, matrix products, generic array operations…): you may want to take a look on their [dedicated page](https://it.mathworks.com/discovery/matlab-multicore.html).

I’m not familiar with matlab internals, but mentioning only linear algebra operations and fft makes me think threading comes only from the underlying libraries used and isn’t a core feature of the language. Is that the case? The documentation page you linked doesn’t seem to be very clear about this.

---

<div class="post-metadata">

**Author:** ![martin.d.maas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/martin.d.maas/32/50964_2.png) [@martin.d.maas](https://discourse.julialang.org/u/martin.d.maas)\
**Post date:** [October 5, 2022, 12:47am UTC](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039/21 "2022-10-05T00:47:08Z")

</div>

Yes, the comparison with C++ is about the random number generator.

Actually, I think that if one forces Julia to use Mersenne Twister it ends up being slower than the C++ code.

```julia
rng = MersenneTwister(1234)

function simple_monte_carlo(n, T)
	x = zeros(n)	
	for i in eachindex(x)
		for t in 1:T	
			x[i] += Random.randn(rng)	
		end	
	end
	return x
end

```

Alternatively, one could swap MersenneTwister for Xoshiro256++ in C++ and report the results in the GitHub repo.

As for a threading efficiency benchmark vs OpenMP, I did run a benchmark a while ago and was very positively surprised by Julia actually being able to benefit from hyperthreading and obtaining an almost 5X vs a single-threaded version in a 4-core PC, something I had never seen with OpenMP where at most I use to get ~3.6X. [Testing Julia blog post](https://www.matecdev.com/posts/numpy-julia-fortran.html).

What has been your experience? Do you have a MWE of where Julia threading falls short? Maybe that could be related to the fact that OpenMP has partial support for SIMD, and in Julia that is implemented on a separate lib (LoopVectorization)?

[Next page](https://discourse.julialang.org/t/blog-using-julia-on-the-hpc/88039.md?page=2)
