# Julia significantly slower than Matlab for simple sin(x) loop – am I doing something wrong?

**URL:** <https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317>\
**Category:** New to Julia\
**Tags:** question, performance\
**Created:** [October 20, 2025, 10:02pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317 "2025-10-20T22:02:41Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![GateOfShadows](https://avatars.discourse-cdn.com/v4/letter/g/dc4da7/32.png) [@GateOfShadows](https://discourse.julialang.org/u/GateOfShadows)\
**Post date:** [October 20, 2025, 10:02pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/1 "2025-10-20T22:02:41Z")

</div>

Hi everyone,

I recently did some benchmark testing between Julia and Matlab, focusing on basic for-loops and vectorized numerical operations. To my surprise, Julia performed significantly worse than Matlab in some very simple scenarios. I’m trying to determine:

- Is this due to my code implementation (missing flags, wrong usage, etc)?
- Or is this an actual limitation of Julia’s current handling of such math-heavy element-wise loops?

**Here’s the test I ran in Julia:**

```julia
using BenchmarkTools

arr = randn(Int(1e7))
vals = similar(arr)

# for loop
@btime @inbounds @simd for i in eachindex(arr)
    vals[i] = sin(arr[i])
end

# for loop with multi-threading (JULIA_NUM_THREADS = 8)
println("threads: $(Threads.nthreads())")
@btime Threads.@threads for i in eachindex(arr)
    vals[i] = sin(arr[i])
end

# vectorized
@btime vals2 = sin.(arr);

```

The outputs are:

- for loop: `922.618 ms (39998981 allocations: 610.34 MiB)`
- for loop with multi-threading: `204.304 ms (39999023 allocations: 610.34 MiB)`
- vectorized: `87.596 ms (4 allocations: 76.29 MiB)`

**Matlab version:**

```matlab
clear;
clc;

arr = randn(1e7,1);
vals = zeros(size(arr));

tic;
for ii = 1:length(arr)
    vals(ii) = sin(arr(ii));
end
t1 = toc;

tic;
vals2 = sin(arr);
t2 = toc;

```

The outputs are

- for loop: `t1 = 0.1735 s`
- vectorized: `t2 = 0.0285 s`

I expected Julia’s loop to be competitive, especially with `@inbounds`, `@simd`, and `@threads`. But in my test, even Matlab’s **loop** beat Julia’s multi-threaded version — and **Matlab’s broadcast** was **3x faster** than Julia’s.

I’m evaluating Julia for scientific computing (currently using Matlab heavily), and would appreciate any insights or correction of my assumptions!

Thanks in advance!

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [October 20, 2025, 10:20pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/2 "2025-10-20T22:20:59Z")

</div>

Welcome! Yes, this is an artifact of how you’re benchmarking. Julia — unlike Matlab — is very function-centric. You get significantly better performance by putting your code inside functions and passing the data as function arguments. And the way you’re using `@btime` there is treating `arr` and `vals` like non-constant globals — as though you wrote that loop outside a function.

You can tell `@btime` to treat these like function arguments with a special `$` flag. This is just a special thing for BenchmarkTools.

```julia-auto
julia> @btime @inbounds @simd for i in eachindex(arr)
           vals[i] = sin(arr[i])
       end
  918.916 ms (39998981 allocations: 610.34 MiB)

julia> @btime for i in eachindex($arr)
           $vals[i] = sin($arr[i])
       end
  90.637 ms (0 allocations: 0 bytes)

```

No need for `@inbounds` and `@simd` here.

I’m not sure what Matlab is doing in its vectorized version, but I’d suspect threads. Broadcast doesn’t use threads by default, but it does provide a speedup:

```julia-auto
julia> Threads.nthreads()
6

julia> @btime Threads.@threads for i in eachindex($arr)
           $vals[i] = sin($arr[i])
       end
  16.317 ms (32 allocations: 3.20 KiB)

```

---

<div class="post-metadata">

**Author:** ![GateOfShadows](https://avatars.discourse-cdn.com/v4/letter/g/dc4da7/32.png) [@GateOfShadows](https://discourse.julialang.org/u/GateOfShadows)\
**Post date:** [October 20, 2025, 10:31pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/3 "2025-10-20T22:31:33Z")

</div>

Thank you so much! Julia is indeed faster than Matlab now with the “$” flag.

---

<div class="post-metadata">

**Author:** ![John\_Gibson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/john_gibson/32/5321_2.png) [@John\_Gibson](https://discourse.julialang.org/u/John_Gibson)\
**Post date:** [October 20, 2025, 10:49pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/4 "2025-10-20T22:49:29Z")

</div>

Worthwhile readings on this topic:

> **[Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/)**
>
> Documentation for The Julia Language.

> **[Noteworthy Differences from other Languages · The Julia Language](https://docs.julialang.org/en/v1/manual/noteworthy-differences/)**
>
> Documentation for The Julia Language.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [October 21, 2025, 6:18am UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/5 "2025-10-21T06:18:05Z")

</div>

> [@GateOfShadows](#):
>
> `vals2 = sin(arr);`

Also keep in mind that when you call code like this in Matlab, it is directly running a highly optimized C or Fortran implementation, possibly multithreaded. So it’s not Julia vs Matlab, but Julia vs C/Fortran. You should not normally expect Julia to outperform Matlab in those cases.

---

<div class="post-metadata">

**Author:** ![Vasily\_Pisarev](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vasily_pisarev/32/7929_2.png) [@Vasily\_Pisarev](https://discourse.julialang.org/u/Vasily_Pisarev)\
**Post date:** [October 21, 2025, 1:21pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/6 "2025-10-21T13:21:28Z")

</div>

Keep in mind that adding `$` is not a performance tip. It is only a trick to substitute something into the `@btime` macro by value instead of by name. If anything, I would suggest writing as follows if you don’t want to write a separate function:

```julia-auto
@btime (@inbounds for i in eachindex(arr)
    vals[i] = sin(arr[i])
end) setup=(arr=$arr, vals=$vals)

```

Typically, you want to benchmark function calls:

```julia-auto
function set_sin!(target, args)
    @inbounds for i in eachindex(target, args)
        target[i] = sin(args[i])
    end
end

@btime set_sin!(vals, arr)

```

As for why Matlab is still faster than single-threaded Julia, there are _two_ possible explanation which may actually come to play together.  
First, as already said. Matlab may use multiple threads by default.  
Second, the algorithm to compute element-wise sine may be tailored for handling array inputs: there might be a faster rearrangement of arithmetic operations if you want to compute multiple sines than just computing each one sequentially.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [October 21, 2025, 2:07pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/7 "2025-10-21T14:07:10Z")

</div>

> [@DNF](#):
>
> Also keep in mind that when you call code like this in Matlab, it is directly running a highly optimized C or Fortran implementation, possibly multithreaded. So it’s not Julia vs Matlab, but Julia vs C/Fortran. You should not normally expect Julia to outperform Matlab in those cases.

As I wrote in another thread, comparing to optimized library routines (like a vectorized `sin(array)` call) is not particularly interesting except to verify that Julia can produce good compiled code. Where you hope to get significant performance improvements is in cases where you have to write _new_ performance-critical code, not simply call one or two library functions.

> [@Julia's applicable context is getting narrower over time?](https://discourse.julialang.org/t/julias-applicable-context-is-getting-narrower-over-time/55042/40):
>
> If all the high-performance code you will ever need has already been written, then you don’t have as much need for a new language. Julia is attractive for people who need to write **new** high-performance code to solve **new** problems, which don’t fit neatly into the boxes provided by existing numpy [or Matlab] library functions.
> 
> It’s not that Julia has some secret sauce that allows it to beat C — it is just that its compiled performance is comparable to that of C (boiling down to the same LLVM backend), so depending on the programmer’s cleverness and time it will sometimes beat C libraries and sometimes not.

---

<div class="post-metadata">

**Author:** ![Mason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mason/32/2423_2.png) [@Mason](https://discourse.julialang.org/u/Mason)\
**Post date:** [October 21, 2025, 2:32pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/8 "2025-10-21T14:32:43Z")

</div>

> [@mbauman](#):
>
> You can tell `@btime` to treat these like function arguments with a special `$` flag. This is just a special thing for BenchmarkTools.
> 
> ```julia-auto
> julia> @btime @inbounds @simd for i in eachindex(arr)
> vals[i] = sin(arr[i])
> end
> 918.916 ms (39998981 allocations: 610.34 MiB)
> 
> julia> @btime for i in eachindex($arr)
> $vals[i] = sin($arr[i])
> end
> 90.637 ms (0 allocations: 0 bytes)
> 
> ```

Just to tack onto this answer, one might wonder why there is a massive speedup here when you interpolate, but not with the broadcast example:

```julia-auto
julia> @btime vals .= sin.(arr);
  85.849 ms (1 allocation: 16 bytes)

julia> @btime $vals .= sin.($arr);
  86.514 ms (0 allocations: 0 bytes)

```

The reason is that in the loop example, julia needed to access an untyped global variable at each iteration of the loop, so every time it loops it has no idea what `vals` or `arr` could possibly be, whereas in the broadcast example, it doesn’t know what `vals` and `arr` are only at the very start before they get passed into the broadcast machinery, at which point it is able to loop knowing what they are.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [October 28, 2025, 10:00pm UTC](https://discourse.julialang.org/t/julia-significantly-slower-than-matlab-for-simple-sin-x-loop-am-i-doing-something-wrong/133317/9 "2025-10-28T22:00:46Z")

</div>

If you’re bottlenecked by these kinds of streaming workloads (you probably aren’t), you could try

```julia
julia> @btime $vals .= sin.($arr);
  82.147 ms (0 allocations: 0 bytes)

julia> using LoopVectorization

julia> @btime @turbo $vals .= sin.($arr);
  13.840 ms (0 allocations: 0 bytes)

julia> @btime @tturbo $vals .= sin.($arr);
  3.332 ms (0 allocations: 0 bytes)

julia> @btime vmapnt!(sin, $vals, $arr);
  11.214 ms (0 allocations: 0 bytes)

julia> @btime vmapntt!(sin, $vals, $arr);
  2.244 ms (0 allocations: 0 bytes)

julia> versioninfo()
Julia Version 1.12.1
Commit ba1e628ee49 (2025-10-17 13:02 UTC)
Build Info:
  Official https://julialang.org release
Platform Info:
  OS: Linux (x86_64-linux-gnu)
  CPU: 36 × Intel(R) Core(TM) i9-10980XE CPU @ 3.00GHz
  WORD_SIZE: 64
  LLVM: libLLVM-18.1.7 (ORCJIT, cascadelake)
  GC: Built with stock GC
Threads: 36 default, 1 interactive, 36 GC (on 36 virtual cores)

```
