# Benchmarking Julia vs. Python vs. R with PyCall and RCall

**URL:** <https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308>\
**Category:** Performance\
**Tags:** benchmarktools\
**Created:** [April 10, 2020, 1:08am UTC](https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308 "2020-04-10T01:08:00Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [April 10, 2020, 1:08am UTC](https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308/1 "2020-04-10T01:08:00Z")

</div>

I’m putting together an Intro to Julia Jupyter notebook and in my section on _Why you should learn Julia_ I’m emphasizing Julia’s performance. This notebook is for colleagues of mine who primarily use R and a few that use Python. I’m including a few lines of simple benchmarks that look like this:

```julia
using BenchmarkTools
using PyCall
using RCall

a = rand(10^7)

@btime pybuiltin("sum")(a)
@btime R"sum($a)"
@btime sum(a)

```

When you execute this code, Julia absolutely obliterates Python and dramatically outperforms R as well. However, the question that I will undoubtedly get is, “How can I be sure that it doesn’t take longer to execute Python/R code via the PyCall/RCall packages?” or something along those lines. People will obviously want to know if this is a fair way to compare speeds and I simply don’t know enough about the way these packages work to answer those questions.

Does anyone here know if this is a fair comparison or if there are indeed additional processes taking place given that `a` was instantiated in Julia but is being operated on in the other language (or for some other reason)? Aside from telling them to measure the speeds themselves in their normal working environments, is there a good way to convince a skeptical crowd that these are legitimate comparisons?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [April 10, 2020, 2:17am UTC](https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308/2 "2020-04-10T02:17:04Z")

</div>

U r passing data back and forth woth python and R

For fair comparison. Do the sum in R not via julia

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [April 10, 2020, 3:09am UTC](https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308/3 "2020-04-10T03:09:06Z")

</div>

On my machine:

### Julia

```julia
julia> using BenchmarkTools

julia> a = rand(10^7);

julia> @benchmark sum($a)
BenchmarkTools.Trial: 
  memory estimate: 0 bytes
  allocs estimate: 0
  --------------
  minimum time: 3.706 ms (0.00% GC)
  median time: 4.215 ms (0.00% GC)
  mean time: 4.229 ms (0.00% GC)
  maximum time: 5.408 ms (0.00% GC)
  --------------
  samples: 1180
  evals/sample: 1

```

### R

```r
> library(microbenchmark)
> a <- runif(1e7)
> microbenchmark(sum(a))
Unit: milliseconds
   expr min lq mean median uq max neval
 sum(a) 8.633446 8.64609 8.781826 8.700741 8.792563 10.75872 100

```

### Python

```python
In [10]: import numpy as np

In [11]: np_a = np.random.rand(10**7)

In [12]: a = np_a.tolist()

In [13]: %timeit np.sum(np_a)
4.08 ms ± 23.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)

In [14]: %timeit sum(a)
35.8 ms ± 157 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)

```

Performance is nice, but it’s not the primary reason I use Julia. Multiple dispatch, the design of the type system, the support for functional programming, and the ecosystem of numerical and scientific computing packages are what draw me to the language.

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [April 10, 2020, 7:09am UTC](https://discourse.julialang.org/t/benchmarking-julia-vs-python-vs-r-with-pycall-and-rcall/37308/4 "2020-04-10T07:09:58Z")

</div>

Converting Julia Arrays to Python Lists takes some time because they have a completely different memory structure. However, passing Julia Arrays as Numpy Arrays is usually very fast.  
The PyCall overhead is in my experience \<\<1ms if no significant amount of data is transferred. To be on the safe side, I suggest to cross-check the Python and R benchmarks using native Python/ R notebooks.  
I did a comparison of Julia to Python for DataFrames, maybe this is useful for you:

> [@Upload of DataFrame data to databases](https://discourse.julialang.org/t/upload-of-dataframe-data-to-databases/36950):
>
> Hi, for evaluation of a possible projet usage I did a comparison of DataFrames.jl to Pandas, with side-by-side examples and timings: Overall, DataFrames.jl performs very well in my experiments, great work! One functionality I could not find out-of-the-box is for writing the content of a DataFrame to a database (e.g. PostgreSQL), analogue to Pandas df.to\_sql(). A simple implementation of the database upload would be (taken mostly from LibPQ.jl documentation): using DataFrames using LibPQ u…
