# Julia Beginner (from Python): Numba outperforms Julia in rewrite. Any tips to improve performance?

**URL:** <https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414>\
**Category:** Performance\
**Tags:** benchmark, python, tullio, loopvectorization\
**Created:** [August 15, 2021, 12:00am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414 "2021-08-15T00:00:50Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 12:00am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/1 "2021-08-15T00:00:50Z")

</div>

Hey! I’m new to Julia (coming from Python) and was trying to benchmark Julia against a part of the apricot python package (submodular optimization). Unfortunately the Python version (using numba) is still a bit faster (1.5x). I read through the Performance optimization chapter and applied the following already:

- put core components into functions,
- used views for array slices,
- using Threads for parallel for loops,
- using fastmath,
- using column first orientation.  
I asked in the Julia Slack and based on someones recommendation we improved the code by 10x but the key part (the calculate\_gains function) is still slower than numba.  
Here’s the code: [https://github.com/rmeinl/apricot-julia](https://github.com/rmeinl/apricot-julia) (in notebooks/ you’ll find python\_reimpl and julia\_reimpl ). Was wondering if anyone had any recommendations on what I could further improve. 🙂

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 15, 2021, 12:09am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/2 "2021-08-15T00:09:32Z")

</div>

> [@rmeinl](#):
>
> the calculate\_gains function

Is it possible for you to post a working example of that function alone?

Edit:

```julia
function get_gains!(X, current_values, idxs, gains)
    @inbounds Threads.@threads for i in eachindex(idxs)
        s = 0.0
        for j in eachindex(current_values)
            s += @fastmath sqrt(current_values[j] + X[j, idxs[i]])
        end
        gains[i] = s
    end
end;

```

If that is the bottleneck, have you tried to use `@tturbo` in the inner (or outer) loops instead `@threads`? Or `@turbo` in the inner loop at least? Seems to me that what is being computed is too simple for standard threading to be useful. If the number of elements of current\_values is large this is also probably an easy use case for a GPU accelerated map reduce.

Edit 2:

```julia
In [25]:
opt2 = LazyGreedy(X_digits);
res2 = @btime fit(opt2, k);
bench2 = @benchmark fit(opt2, k);
t2 = bench2.times[1]/1e9;
  13.836 s (22202 allocations: 8.95 GiB)

```

Those allocations, unless you know exactly where are they coming from, are an indication of a type instability. I find very hard to debug a notebook, but that may be an indication that the program can be much faster.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [August 15, 2021, 2:27am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/3 "2021-08-15T02:27:40Z")

</div>

FYI `@inbounds Threads.@threads for` does not remove the bound check in the loop body. See: [What is the cause of this performance difference between Julia and Cython? - #4 by tkf](https://discourse.julialang.org/t/what-is-the-cause-of-this-performance-difference-between-julia-and-cython/58886/4)

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 3:32am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/4 "2021-08-15T03:32:22Z")

</div>

Thanks for the pointer! I put it inside the loop but it didn’t give any performance improvement

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 3:41am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/5 "2021-08-15T03:41:28Z")

</div>

I tried using @tturbo instead of @Threads and it was slightly slower. I also tried using @avx and @turbo in the inner loop and it threw this error:

```julia
TaskFailedException

    nested task error: MethodError: no method matching subsetview(::VectorizationBase.FastRange{Int64, Static.StaticInt{0}, Static.StaticInt{1}, Int64}, ::Static.StaticInt{1}, ::Int64)
    Closest candidates are:
      subsetview(::VectorizationBase.StridedPointer{T, N, C, B, R, X, O}, ::Static.StaticInt{I}, ::Integer) where {T, N, C, B, R, X, O, I} at /Users/ricomeinl/.julia/packages/LoopVectorization/yPPLg/src/vectorizationbase_compat/subsetview.jl:2

```

> [@lmiq](#):
>
> Is it possible for you to post a working example of that function alone?

Sure, I created a little script that only uses that function. Its here: [https://github.com/rmeinl/apricot-julia/blob/main/src/calc\_gains.jl](https://github.com/rmeinl/apricot-julia/blob/main/src/calc_gains.jl)

> [@lmiq](#):
>
> Those allocations, unless you know exactly where are they coming from, are an indication of a type instability.

Does that mean if I explicitly add type annotations to the function I could speed it up?

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [August 15, 2021, 4:14am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/6 "2021-08-15T04:14:45Z")

</div>

What size is the data, i.e. what are `d` and `n` in the script you link? (Without having to solve ScikitLearn installation problems… the times would work as well with `rand(d,n)`, I presume.)

And is `get_gains!` meant only to work with `idxs == 1:n`, or will these in general be in some other order? That might make a difference, and would certainly allow simplification if not required.

And finally, does it matter that X is transposed? Your loop (if I’m reading it right) would prefer that `j` indexes adjacent elements, which would I think be true with `X_digits = permutedims(abs.(digits_data["data"]));`.

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 4:44am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/7 "2021-08-15T04:44:16Z")

</div>

> [@mcabbott](#):
>
> What size is the data, i.e. what are `d` and `n` in the script you link?

(d, n) = (54, 581012). I changed the script to use a random matrix.

> [@mcabbott](#):
>
> And is `get_gains!` meant only to work with `idxs == 1:n` , or will these in general be in some other order?

`idxs` is simply an array of indices, they could be in any order.

> [@mcabbott](#):
>
> And finally, does it matter that X is transposed?

The reason I transposed it from (number of elements, dimensions) to (dimensions, number of elements) is because I read that Julia uses matrices in column first orientation, which would give me a speedup if the matrix was transposed.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [August 15, 2021, 5:20am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/8 "2021-08-15T05:20:16Z")

</div>

OK, great. Here are times I get, with this size. I think the transpose might not be doing quite what you’re thinking, but especially for random (or shuffled) indices it will matter quite a bit. (Both here and in python.) In that case, fancy packages seem to get me almost another factor of 2.

```julia
Xmat = rand(54, 581_000); # approx true size now
Xlazy = transpose(permutedims(Xmat));
Xmat == Xlazy # same numbers, but different memory: strides(Xmat), strides(Xlazy)

ind_order = collect(1:581_000);
ind_rand = rand(1:581_000, 581_000);

current = rand(54);
gains = zeros(581_000);
# get_gains_0!(Xmat, current, ind_rand, gains) # version with bounds checks

@btime get_gains!($Xmat, $current, $ind_order, $gains); # 5.303 ms
@btime get_gains!($Xmat, $current, $ind_rand, $gains); # 21.661 ms
@btime get_gains!($Xlazy, $current, $ind_order, $gains); # 8.233 ms
@btime get_gains!($Xlazy, $current, $ind_rand, $gains); # 87.751 ms

using Tullio
get_7!(X, cv, idxs, gains) = @tullio gains[i] = sqrt(cv[j] + X[j, idxs[i]]) avx=false
using LoopVectorization
get_8!(X, cv, idxs, gains) = @tullio gains[i] = sqrt(cv[j] + X[j, idxs[i]])
# Assuming this is the relevant case:
@btime get_7!($Xmat, $current, $ind_rand, $gains); # 13.965 ms
@btime get_8!($Xmat, $current, $ind_rand, $gains); # 11.944 ms

# Eager permutedims is quite expensive, but perhaps only once:
@btime permutedims($Xmat); # 199.528 ms 
@btime permutedims($(permutedims(Xmat))); # 61.824 ms

```

---

<div class="post-metadata">

**Author:** ![genkuroki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/genkuroki/32/18030_2.png) [@genkuroki](https://discourse.julialang.org/u/genkuroki)\
**Post date:** [August 15, 2021, 12:13pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/9 "2021-08-15T12:13:04Z")

</div>

> [@lmiq](#):
>
> ```julia
> 13.836 s (22202 allocations: 8.95 GiB)
> 
> ```
> 
> Those allocations,

Those allocations come from `idxs = findall(==(0), mask)`. (See sixth line from the bottom of In [23] of [https://github.com/rmeinl/apricot-julia/blob/4ab1e5f397155b116a36769e59d446d4f465c40a/notebooks/julia\_benchmark.ipynb](https://github.com/rmeinl/apricot-julia/blob/4ab1e5f397155b116a36769e59d446d4f465c40a/notebooks/julia_benchmark.ipynb))

The `idxs` will be a vector of `Int64` with length of several hundred thousand.

```julia
n = 581012
mask = zeros(Int8, n)
mask[1] = 1
@time idxs = findall(==(0), mask);

```

```julia
0.007981 seconds (21 allocations: 9.001 MiB)

```

These are repeated 1000 times to produce about 2×10⁴ allocations: 9 GiB.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [August 15, 2021, 12:26pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/10 "2021-08-15T12:26:42Z")

</div>

> [@rmeinl](#):
>
> using numba

Numba is very strict, and some times even distort the semantics of python because it can’t tolerate any type instability, for example:

```julia
1 if np.random.rand() < 0.5 else 0.5

```

will be turned into returning `1.0` or `0.5` (always Float).

All I’m trying to say is that if your functions work nicely in Numba naturally, you’re basically comparing a C function to Julia function, assuming they are doing similar things algorithmically, I’d hope neither is much slower than the other!

Btw in your code, `X` is a global non-constant, this causes type instability everywhere.

* * *

The reason to use Julia instead of Numba is of course that Numba only works with a subset of Python semantics and often doesn’t work well with third-party packages that have their own C/C++ backend. (outside of numpy ecosystem)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [August 15, 2021, 12:45pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/11 "2021-08-15T12:45:11Z")

</div>

As I commented in another thread, micro-optimizing a single vectorized function that does a single operation (sums the `sqrt` of array elements in your case) is not really the best way to get performance improvements:

> [@Comparing Numba and Julia for a complex matrix computation](https://discourse.julialang.org/t/comparing-numba-and-julia-for-a-complex-matrix-computation/63703/3):
>
> I should add that this is almost certainly the wrong way to go about getting performance improvements. “Highly optimized” code in Python usually means code that is “vectorized” as much as possible — broken into a series of relatively simple array operations that can individually call fast library operations. (For example, calling atan on a bunch of numbers in your code above.) You can write the same sort of code in Julia, of course, but it won’t be magically faster than Numpy — there’s no s…

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 15, 2021, 1:05pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/12 "2021-08-15T13:05:16Z")

</div>

Following the advice above, maybe the best would be to try to avoid completely the creation of this intermediary array. It may cost as much as the gains evaluation, or even more.

---

<div class="post-metadata">

**Author:** ![genkuroki](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/genkuroki/32/18030_2.png) [@genkuroki](https://discourse.julialang.org/u/genkuroki)\
**Post date:** [August 15, 2021, 1:21pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/13 "2021-08-15T13:21:24Z")

</div>

> [@genkuroki](#):
>
> Those allocations come from `idxs = findall(==(0), mask)` .

I have tried to remove those allocations.

`mask = zeros(Int8, n)`  
 → _deleted_

`idxs = 1:n`  
 → `idxs = collect(1:n)`

`current_values += view(optimizer.X, :, best_idx)`  
 → `current_values .+= view(optimizer.X, :, best_idx)`

`current_concave_values .= sqrt.(current_values)`  
 → _deleted_

`current_concave_values_sum = sum(current_concave_values)`  
 → `current_concave_values_sum = sum(sqrt, current_values)`

`mask[best_idx] = 1; idxs = findall(==(0), mask)`  
 → `popat!(idxs, findfirst(==(best_idx), idxs))`

The last one is essential.

See In [10] of [my Jupyter notebook](https://github.com/genkuroki/public/blob/7c1ab5bce26f9dffd0e494796b00c4602cdb988f/0016/NaiveGreedy.ipynb) for more details.

Verification of the coincidence of the two results:

```julia
@show res1rev.ranking == res1.ranking
@show res1rev.gains ≈ res1.gains;

```

```julia
res1rev.ranking == res1.ranking = true
res1rev.gains ≈ res1.gains = true

```

Subset of `show(to)` of the original version:

```julia
    Tot / % measured: 20.2s / 94.6% 9.19GiB / 96.1%
   calc_gains 1.00k 12.0s 62.9% 12.0ms 13.3MiB 0.15% 13.6KiB
   select_next 1.00k 5.76s 30.2% 5.76ms 8.82GiB 100% 9.03MiB

```

Subset of `show(to)` of the revised version:

```julia
    Tot / % measured: 13.3s / 98.1% 274MiB / 5.31%
   calc_gains 1.00k 11.5s 88.7% 11.5ms 6.10MiB 42.0% 6.25KiB
   select_next 1.00k 184ms 1.41% 184μs 3.93MiB 27.0% 4.02KiB

```

**5.76s, 8.82GiB → 184ms, 3.93MiB**

**Postscript:** Although not essential in this case, it should be noted that, in general, `sum(f.(A))` for an array `A` will produce an allocation for `f.(A)`. If `sum(f, A)` is used instead, the allocation will be zero; `mean(f.(A))`, `minimum(f.(A))`, etc. should likewise be rewritten as `mean(f, A)`, `maximum(f, A)`, etc.

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 4:36pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/14 "2021-08-15T16:36:28Z")

</div>

Very nice! Thank you. This slashed the time of both my NaiveGreedy and LazyGreedy by almost 2x.

```julia
# NaiveGreedy old
18.885 s (104754 allocations: 8.80 GiB)

# NaiveGreedy new
10.610 s (83658 allocations: 15.04 MiB)

# LazyGreedy old
13.836 s (22202 allocations: 8.95 GiB)

# LazyGreedy new
6.264 s (200 allocations: 167.93 MiB)

```

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 4:38pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/15 "2021-08-15T16:38:23Z")

</div>

Fair point. I was just wondering why the numba version was still faster. I’d have expected them to be the same and then the Julia “glue” code to be a lot faster

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 15, 2021, 4:42pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/16 "2021-08-15T16:42:45Z")

</div>

Not sure if I understand. Do you recommend using permutedims over transpose? And is tullio faster than fastmath? How would you apply your recommendations to this script: [https://github.com/rmeinl/apricot-julia/blob/main/src/calc\_gains.jl](https://github.com/rmeinl/apricot-julia/blob/main/src/calc_gains.jl)

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 15, 2021, 4:58pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/17 "2021-08-15T16:58:51Z")

</div>

`transpose` is lazy, so you don’t get the memory benefits of looping over columns just because you use `transpose`.

Better to use `permutedims` from the get-go. Try and read in the data to have a “long” shape from the very beginning.

It would be great if you could decompose changes you made into two parts

1. Reducing allocations, making sure to loop along contiguous memory etc.
2. The addition of `@tullio`, `@fastmath` etc.

I would bet that you can get Numba-level performance with juse 1. and then 2. should make it go past Numba, but I’m not sure.

---

<div class="post-metadata">

**Author:** ![jakobnissen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jakobnissen/32/13477_2.png) [@jakobnissen](https://discourse.julialang.org/u/jakobnissen)\
**Post date:** [August 16, 2021, 11:39am UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/18 "2021-08-16T11:39:03Z")

</div>

I understand that you’re just trying to explain best practices, but having actually benchmarked this problem, you are not correct.

1. The fact that transpose is lazy doesn’t matter here, because the processing is done on a type wrapping `Matrix{Float64}`, so `convert(Matrix{Float64}, A)` is automatically called.
2. The large majority of time is spent in a single function. Removing allocations does have a large effect, but nowhere enough to get it to Numba speed.
3. Numba is actually really fast here. Like, really fast. Fast enough that I seriously doubt you can beat it using only Base Julia.

Without having played around with Tullio etc, I believe the major reason Numba is faster than Base Julia here is that it schedules its threads more effectively than `Threads.@threads` for the hot inner loop.

All this is, of course, without thinking about the algorithm itself. It may very well be that @stevengj is on point in his comment, and this problem can be solved more effectively by doing more scalar operations instead of boradcasted operations.

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [August 16, 2021, 12:44pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/19 "2021-08-16T12:44:17Z")

</div>

Edit: I splitted the python and Julia code into separate files (the original one, not the one optimized by @genkuroki). This is what I get:

```julia
% python3 apricot.py
-------------------------------------------------
Fitting time: 12.442765712738037
125818832 9897.307670089194
-------------------------------------------------

```

```julia
% julia -t auto -i apricot.jl
  Activating environment at `~/Drive/Work/JuliaPlay/apricot/Project.toml`
---------------------------------------------
 Fitting time:
 12.268594 seconds (22.16 k allocations: 8.951 GiB, 0.22% gc time)
125819832 9897.307670089054
---------------------------------------------

```

The files are here: [https://gist.github.com/lmiq/828133335a83088e63f21152416c7f7c](https://gist.github.com/lmiq/828133335a83088e63f21152416c7f7c)

The “hot loop” takes 2.5 seconds of those:

```julia
 ───────────────────────────────────────────────────────────────────
                            Time Allocations      
                    ────────────────────── ───────────────────────
  Tot / % measured: 53.3s / 4.86% 19.9GiB / 0.00%    

 Section ncalls time %tot avg alloc %tot avg
 ───────────────────────────────────────────────────────────────────
 hot loop 9.43M 2.59s 100% 275ns 0.00B - % 0.00B
 ───────────────────────────────────────────────────────────────────

```

Thus, I am not sure if I’m doing everything wrong, or if we are not really comparing apples to apples in general.

---

<div class="post-metadata">

**Author:** ![rmeinl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rmeinl/32/28163_2.png) [@rmeinl](https://discourse.julialang.org/u/rmeinl)\
**Post date:** [August 16, 2021, 2:14pm UTC](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414/20 "2021-08-16T14:14:26Z")

</div>

@lmiq you’re comparing Julia’s `LazyGreedy` with Python’s `NaiveGreedy`. I should’ve made it more explicit but the Python `fit()` method refers to the `NaiveGreedy` which is a lot faster in Julia.

[Next page](https://discourse.julialang.org/t/julia-beginner-from-python-numba-outperforms-julia-in-rewrite-any-tips-to-improve-performance/66414.md?page=2)
