# Trying to speed-up Laplace Equation Solver

**URL:** <https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862>\
**Category:** Performance\
**Created:** [July 29, 2020, 9:25am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862 "2020-07-29T09:25:24Z")\
**Posts on this page:** 17\
**Page:** 2

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 29, 2020, 6:05pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/21 "2020-07-29T18:05:55Z")

</div>

@stevengj I agree its fast enough for practical purposes. I just wanted to be sure that I have tried out all possible options and achieving the maximum speedup I can get. Its just for my practice with the language.

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 29, 2020, 6:08pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/22 "2020-07-29T18:08:43Z")

</div>

> [@PetrKryslUCSD](#):
>
> Check out the Julia code in this benchmark: [Interesting NASA “Language-Comparison” repository](https://discourse.julialang.org/t/interesting-nasa-language-comparison-repository/31356/26)

@PetrKryslUCSD This is a great link you directed me to! In fact, I was trying this with NVIDIA codes for C and Fortran.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [July 29, 2020, 6:18pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/23 "2020-07-29T18:18:00Z")

</div>

Just curious: Why is the registered LoopVectorization stuck at 0.6.30?

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 29, 2020, 6:48pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/24 "2020-07-29T18:48:45Z")

</div>

> [@stillyslalom](#):
>
> You can attain a significant speedup by moving the `@fastmath` annotation to the outer loop to allow it to operate on the whole loop body, and it’s faster yet with LoopVectorization.jl’s `@avx` macro

I tried these improvements in three places: Windows 10 System, Ubuntu 20.04 WSL-2 in the same Windows 10 System and Ubuntu 16.04 Server. Sepcifically, I compared using `@inbounds @fastmath` and using `@avx`.

I was able to get improvements in Windows 10 as expected. My code ran 25% faster after using `@avx` macro. However, surprisingly, I could not find any difference whatsoever in the WSL-2 Ubuntu or the Ubuntu server. How is this possible? (I may be wrong, but I have this feeling that Windows 10 is able to leverage these benefits out of the Intel core but Ubuntu is unable to).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 30, 2020, 1:32pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/25 "2020-07-30T13:32:10Z")

</div>

> [@PetrKryslUCSD](#):
>
> Just curious: Why is the registered LoopVectorization stuck at 0.6.30?

Because FinEtools places an upper bound on LoopVectorization at “0.6”. Without this upper bound, the package manager can select 0.8.20.

> [@SambitMishra98](#):
>
> I could not find any difference whatsoever in the WSL-2 Ubuntu

Could you provide `versioninfo()` for all three?  
As this was the same system as the Windows benchmarks, what were the relative timings between them?

Also, you could try

```julia
function calcNext2!(Anew, A)
    error = zero(eltype(A))
    m, n = size(A)
    @avx unroll=(1,1) for j = 2:(m-1)
        for i = 2:(n-1)
            Anew[i,j] = 0.25*( A[i+1, j]
                             + A[i-1, j]
                             + A[i ,j-1]
                             + A[i ,j+1])
            error = max(error, abs(Anew[i,j] - A[i,j]))
        end
    end
    return error
end

```

to prevent unrolling the loops. This seems to help performance slightly.  
It increases the number of instructions needed to evaluate the loop by about 30%, but it decreases the number of L1d misses, which in turn increases the instructions/clock.

What I said earlier about the `@pstats` macro (which you can try using LinuxPerf on the Ubuntu server (assuming `/proc/sys/kernel/perf_event_paranoid` is \<= 1, or that you have permission to `sudo sh -c "echo 1 > /proc/sys/kernel/perf_event_paranoid"`) is that when you use `@avx`, the CPU has to do a small fraction of the amount of the amount of work as without it. But that the CPU then goes from being very busy, to spending most of its time idle, not doing anything while waiting for memory to arrive.

I suspect that the hardware prefetcher doesn’t know what to do with the memory access pattern when we don’t unroll, and either gives up or (worse) prefetches data we don’t need. But I’m no expert when it comes to this.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [July 30, 2020, 3:08pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/26 "2020-07-30T15:08:48Z")

</div>

Of course, I forgot that when I do “update” it will not tell me that there is a package update  
that is possible.

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [July 30, 2020, 4:21pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/27 "2020-07-30T16:21:50Z")

</div>

This is why I originally considered the idea of forcing all releases to have upper bounds to be controversial (since there is no good communication for updates). A possible aid with this is to subscribe to the releases of dependencies on GitHub.

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 30, 2020, 6:19pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/28 "2020-07-30T18:19:35Z")

</div>

> [@Elrod](#):
>
> Could you provide `versioninfo()` for all three?

WSL-2 Ubuntu 20.04 in Windows 10:

```julia
Julia Version 1.4.2
Commit 44fa15b150* (2020-05-23 18:35 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Core(TM) i7-8550U CPU @ 1.80GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-8.0.1 (ORCJIT, skylake)
Environment:
  JULIA_PKG_SERVER = pkg.juliacomputing.com
  JULIA_DEPOT_PATH = /home/sambit98/.juliapro/JuliaPro_v1.4.2-1:/mnt/d/SAMBIT/JULIA/JuliaInstallation/Jn/JuliaPro-1.4.2-1/Julia/local/share/julia:/mnt/d/SAMBIT/JULIA/JuliaInstallation/JuliaPro-1.4.2-1/Ju/shlia/share/julia

```

Ubuntu 16.04 Server:

```julia
Julia Version 1.4.2
Commit 44fa15b150* (2020-05-23 18:35 UTC)
Platform Info:
  OS: Linux (x86_64-pc-linux-gnu)
  CPU: Intel(R) Core(TM) i7-7700 CPU @ 3.60GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-8.0.1 (ORCJIT, skylake)

```

Windows 10:

```julia
Julia Version 1.4.2
Commit 44fa15b150* (2020-05-23 18:35 UTC)      
Platform Info:
  OS: Windows (x86_64-w64-mingw32)
  CPU: Intel(R) Core(TM) i7-8550U CPU @ 1.80GHz
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-8.0.1 (ORCJIT, skylake)
Environment:
  JULIA_EDITOR = "C:\Users\Ashok Kumar Mishra\AppData\Local\atom\app-1.49.0\atom.exe" -a
  JULIA_NUM_THREADS = 4

```

> [@Elrod](#):
>
> As this was the same system as the Windows benchmarks, what were the relative timings between them?

The following are the times taken. I ran the code 2-3 times and chose the time taken by the last execution in each of the cases. I ran Ubuntu Julia in Terminal and Windows 10 Julia in Juno, and tried to keep all factors as consistent as possible.

Using `@inbounds @fastmath`  
WSL-2 Ubuntu 20.04 : 38.616142 seconds (674.78 k allocations: 31.788 MiB, 0.07% gc time)  
Server Ubuntu 16.04 : 34.896664 seconds (674.78 k allocations: 31.786 MiB, 0.05% gc time)  
Windows 10 : 39.097357 seconds (672.78 k allocations: 31.679 MiB)

Using `@avx`  
WSL-2 Ubuntu 20.04 : 42.330296 seconds (14.95 M allocations: 742.403 MiB, 0.48% gc time)  
Server Ubuntu 16.04 : 35.750165 seconds (13.52 M allocations: 660.662 MiB, 0.34% gc time)  
Windows 10 : 38.948415 seconds (13.47 M allocations: 658.412 MiB, 0.71% gc time)

I realised that I incorrectly came to the conclusion that using `@avx` shows a significant difference in Windows 10 alone, since I was using Revise to change the function content and execute (which had given me the 25% speedup). So in the end, based on these executions, `@avx` is slightly decreasing the speed.

Using `@avx unroll(1,1)`  
WSL-2 Ubuntu 20.04: 39.749676 seconds (14.23 M allocations: 706.557 MiB, 0.47% gc time)  
Server Ubuntu 16.04: 30.470257 seconds (12.75 M allocations: 622.198 MiB, 0.40% gc time)  
Windows 10: 36.784031 seconds (12.71 M allocations: 619.726 MiB, 0.73% gc time)

So the last suggestion is the best, although it’s still slower than the Fortran code which takes 28 seconds in the Server Ubuntu.

> [@chakravala](#):
>
> This is why I originally considered the idea of forcing all releases to have upper bounds to be controversial (since there is no good communication for updates). A possible aid with this is to subscribe to the releases of dependencies on GitHub.

Based on [https://www.youtube.com/watch?v=xPhnJCAkI4k](https://www.youtube.com/watch?v=xPhnJCAkI4k) I think these issues will not be there by v1.5. I may be wrong since I didn’t understand much from the video as a noob.

---

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [July 30, 2020, 8:23pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/29 "2020-07-30T20:23:34Z")

</div>

That’s a lot of allocations. Are your matrices non-const globals? Also, if you run to convergence, you may expect a runtime difference due to different number of iterations between implementations.

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 31, 2020, 2:08am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/30 "2020-07-31T02:08:55Z")

</div>

> [@Vasily\_Pisarev](#):
>
> Are your matrices non-const globals?

Matrices are not globals. I have initialised them and used them as shown below.

```julia
push!(LOAD_PATH,".");
using laplace2d

const n = 4096;
const m = 4096;

A,Anew = initialize(m,n);
println("Jacobi relaxation Calculation: ",n,"x",m," mesh");
@time Iterate!(A,Anew,m,n);

```

Wherein, I used the following function:

```julia
    function initialize(m,n)
        A = zeros(Float64,n,m);
        A[1,:] = ones(Float64,n );
        Anew = deepcopy(A);
        return(A,Anew)
    end

```

> [@Vasily\_Pisarev](#):
>
> Also, if you run to convergence, you may expect a runtime difference due to different number of iterations between implementations.

A fixed number of 1000 iterations is executed all the time.

Although I have a feeling that the [definition of arrays can be modified to my needs](https://www.youtube.com/watch?v=jS9eouMJf_Y&t=632s), I don’t know how I should do that (or if there is any scope of improvement in that direction).

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 31, 2020, 6:41am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/31 "2020-07-31T06:41:56Z")

</div>

> [@SambitMishra98](#):
>
> Using `@inbounds @fastmath`  
> WSL-2 Ubuntu 20.04 : 74.78 k allocations: 31.788 MiB, 0.07% gc time  
> Server Ubuntu 16.04 : 674.78 k allocations: 31.786 MiB, 0.05% gc time  
> Windows 10 : 672.78 k allocations: 31.679 MiB
> 
> Using `@avx`  
> WSL-2 Ubuntu 20.04 : 14.95 M allocations: 742.403 MiB, 0.48% gc time  
> Server Ubuntu 16.04 : 13.52 M allocations: 660.662 MiB, 0.34% gc time  
> Windows 10 : 13.47 M allocations: 658.412 MiB, 0.71% gc time
> 
> Using `@avx unroll(1,1)`  
> WSL-2 Ubuntu 20.04: 14.23 M allocations: 706.557 MiB, 0.47% gc time  
> Server Ubuntu 16.04: 12.75 M allocations: 622.198 MiB, 0.40% gc time  
> Windows 10: 12.71 M allocations: 619.726 MiB, 0.73% gc time

And, why are there so many more allocations with the `@avx` versions than the `@inbounds @fastmath` versions?

Is this counting counting compilation?

As for `initialize`, do you really need `Anew` to be a copy of `A`? In the `calcNext` function, you’re overwriting its contents without reading. You also don’t need to make a vector of ones:

```julia
function initialize(m,n)
    A = zeros(Float64,n,m);
    A[1,:] .= 1
    Anew = similar(A);
    return(A,Anew)
end

```

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 31, 2020, 9:58am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/32 "2020-07-31T09:58:26Z")

</div>

> [@Elrod](#):
>
> As for `initialize` , do you really need `Anew` to be a copy of `A` ?

I did not pay much attention to initializing since the bottleneck of the code was `calc!` and `copy!` functions. I am comparing my results with Fortran code. The error in that code is calculated by considering the edges of the domain too, which are not modified throughout the simulation. Hence, using `similar` is not preferred here. I have incorporated the other changes.

> [@Elrod](#):
>
> Is this counting counting compilation?

I am not sure about this, can you tell me how to find out? I ran 2 simulations consecutively with `using Revise` and got the following results.  
RUN 1: `36.588347 seconds (12.55 M allocations: 612.136 MiB, 0.32% gc time)`  
RUN 2: `30.291741 seconds (174 allocations: 8.188 KiB)`

Also it seems like this is getting more involved, so I have included the entire code below. I have jacobi.jl as the main program and laplace2d.jl code written like a module.

jacobi.jl

```julia
push!(LOAD_PATH,".");

using laplace2d

const n = 4096;
const m = 4096;

A,Anew = initialize2(m,n);

println("Jacobi relaxation Calculation: ",n,"x",m," mesh");

@time Iterate!(A,Anew,m,n);

```

laplace2d.jl

```julia
module laplace2d

using LoopVectorization

export initialize
export Iterate!

    function initialize(m,n)

        A = zeros(Float64,n,m);
        A[1,:] .= 1

        Anew = deepcopy(A);

        return(A,Anew)
    end

    function calcNext!(Anew,A,m,n)
        println(Anew)
        error = 0.0::Float64;

        @avx unroll=(1,1) for j=2:(m-1)
            for i = 2:(n-1)

                 Anew[i,j] = 0.25*( A[i+1, j]
                                  + A[i-1, j]
                                  + A[i ,j-1]
                                  + A[i ,j+1] );

                error = max(error, abs( Anew[i,j] - A[i,j] ) );
            end
        end

        return(error);
    end

    function Iterate!(A,Anew,m,n)
        Error = 1.0::Float64;
        iter = 0;

        tol = 1.0e-6::Float64;
        iter_max = 10;
        while (Error>tol && iter<iter_max)

            Error = calcNext!(Anew,A,m,n);
            copyto!(A,Anew);

            if(mod(iter,100) ≈ 0); println(iter,"\t",Error); end

            iter += 1;
        end
    end
end

```

Output from Julia code:

```julia
Jacobi relaxation Calculation: 4096x4096 mesh
0 0.25
100 0.002397062376472081
200 0.0012043203299441085
300 0.000803706300640139
400 0.0006034607481846255
500 0.0004830019181236156
600 0.00040251372422966947
700 0.00034514687472469996
800 0.0003021170179446919
900 0.0002685514561578395
 30.395864 seconds (12.75 M allocations: 622.198 MiB, 0.40% gc time)

```

Output from Fortran code:

```julia
Jacobi relaxation Calculation: 4096 x 4096 mesh
    0 0.25000000000000000000
  100 0.00239706248976290226
  200 0.00120432034600526094
  300 0.00080370629439130425
  400 0.00060346076497808099
  500 0.00048300193157047033
  600 0.00040251371683552861
  700 0.00034514686558395624
  800 0.00030211702687665820
  900 0.00026855146279558539
 completed in 28.006 seconds

```

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 31, 2020, 10:11am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/33 "2020-07-31T10:11:26Z")

</div>

With 4096x4096 matrices, the memory problems are going to be even worse than when they were 512x512.

Another optimization is to drop the copies in the loop:

```julia
function Iterate!(A,Anew,m,n)
    Error = 1.0::Float64;
    iter = 0;

    tol = 1.0e-6::Float64;
    iter_max = 10;
    Aorig = A
    while (Error>tol && iter<iter_max)

        Error = calcNext!(Anew,A,m,n);
        Anew, A = A, Anew

        if(mod(iter,100) ≈ 0); println(iter,"\t",Error); end

        iter += 1;
    end
    A === Aorig || copyto!(Aorig, A)
    nothing
end

```

Why to get rid of any unnecessary `copyto!`s:

```julia
julia> A = rand(4096, 4096); Anew = similar(A);

julia> @btime calcNext3!($Anew, $A)
  18.138 ms (0 allocations: 0 bytes)
 0.9740572038986998

julia> @btime copyto!($Anew, $A);
  13.360 ms (0 allocations: 0 bytes)

```

They’re almost as slow as `calcNext3!` (the `@avx unroll=(1,1)` version).

FWIW, I get:

```julia
julia> A,Anew = initialize(m,n);

julia> @time Iterate!(A,Anew,m,n);
0 0.25
100 0.002397062376472081
200 0.0012043203299441085
300 0.000803706300640139
400 0.0006034607481846255
500 0.0004830019181236156
600 0.00040251372422966947
700 0.00034514687472469996
800 0.0003021170179446919
900 0.0002685514561578395
 18.105712 seconds (213 allocations: 9.109 KiB)

```

This was with `iter_max = 1_000`.

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [July 31, 2020, 11:18am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/34 "2020-07-31T11:18:29Z")

</div>

@Elrod that is just amazing !! I didn’t expect to even go this far! Thanks a lot for supporting me till here! Now the Julia code is 40% faster than the Fortran code!

I have a doubt regarding running the code twice. The following is what I got.

```julia

julia> include("jacobi.jl")
[Info: Precompiling laplace2d [top-level]
Jacobi relaxation Calculation: 4096x4096 mesh
0 0.25
100 0.002397062376472081
200 0.0012043203299441085
300 0.000803706300640139
400 0.0006034607481846255
500 0.0004830019181236156
600 0.00040251372422966947
700 0.00034514687472469996
800 0.0003021170179446919
900 0.0002685514561578395
 24.432541 seconds (12.70 M allocations: 619.979 MiB, 0.99% gc time)

julia> include("jacobi.jl")
Jacobi relaxation Calculation: 4096x4096 mesh
0 0.25
100 0.002397062376472081
200 0.0012043203299441085
300 0.000803706300640139
400 0.0006034607481846255
500 0.0004830019181236156
600 0.00040251372422966947
700 0.00034514687472469996
800 0.0003021170179446919
900 0.0002685514561578395
 17.869136 seconds (173 allocations: 7.859 KiB)

```

The number of allocations is less only when I run the code more than once. I am not clear on why that is so (you had asked if its counting compilation before this, I didn’t get that as well). Why is there so much difference between the two executions? Also, does that mean if I somehow compile the code and run it as an executable, I would get performance akin to the second execution?

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [July 31, 2020, 1:20pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/35 "2020-07-31T13:20:42Z")

</div>

> [@SambitMishra98](#):
>
> Why is there so much difference between the two executions?

The first time you use `@avx`, you need to not only compile that function, but LoopVectorization as well. This will take on the order of 5+ seconds.

> [@SambitMishra98](#):
>
> Also, does that mean if I somehow compile the code and run it as an executable, I would get performance akin to the second execution?

Yes.

---

<div class="post-metadata">

**Author:** ![SambitMishra98](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sambitmishra98/32/14263_2.png) [@SambitMishra98](https://discourse.julialang.org/u/SambitMishra98)\
**Post date:** [August 1, 2020, 2:03am UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/36 "2020-08-01T02:03:33Z")

</div>

I created another code `JacobiSysImage.jl` based on PackageCompiler [documentation](https://julialang.github.io/PackageCompiler.jl/dev/examples/plots/) and [video](https://www.youtube.com/watch?v=d7avhSuK2NA).

Code for `JacobiSysImage.jl`

```julia
using PackageCompiler
create_sysimage(:LoopVectorization,sysimage_path="jacobi_sys_image.so",precompile_execution_file="jacobi.jl") 

```

Then I got a two-line terminal execution of running the main code using compiled packages with the help of a [stackoverflow discussion](https://stackoverflow.com/questions/59806941/how-to-use-julias-packagecompiler-to-build-a-quick-launch-environment-for-plots).

```julia
$ julia JacobiSysImage.jl
$ julia -J jacobi_sys_image.so -q -e 'include("jacobi.jl")' 

```

I decreased the execution time further this way, and now the code is around 50% faster than the compiled Fortran code. I should also mention that `PackageCompiler` took a few minutes to create `jacobi_sys_image.so`.

The whole idea behind this was to execute the main code using two lines of terminal script, similar to what was done in Fortran.

Although there is an option to create an executable app in `PackageCompiler`, I did not find that easy enough to execute or worth the effort (correct me if I am wrong).

Also, is there no way to get allocations minimal with the first execution of the code? I am still getting a large number of allocations when I execute the code (although the performance is similar to when I execute the code the second time).  
`15.010896 seconds (1.29 M allocations: 64.140 MiB, 0.03% gc time)`

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [August 1, 2020, 10:14pm UTC](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862/37 "2020-08-01T22:14:26Z")

</div>

> [@SambitMishra98](#):
>
> Also, is there no way to get allocations minimal with the first execution of the code? I am still getting a large number of allocations when I execute the code (although the performance is similar to when I execute the code the second time).

I believe the functions you defined in the script are still being recompiled, but the library dependencies (in particular, LoopVectorization) don’t have to be recompiled, making this much faster. There’s a risk of some of LoopVectorization needing to get recompiled anyway because it’s currently passing `gensym`ed variables as type parameters. I’ll address that in a future release.

If you create an app, none of it should have to recompile.

[Previous page](https://discourse.julialang.org/t/trying-to-speed-up-laplace-equation-solver/43862.md?page=1)
