# What Is a Simple Example of Julia Solving the Two Language Problem?

**URL:** <https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659>\
**Category:** General Usage\
**Tags:** question, teaching\
**Created:** [April 19, 2023, 2:30pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659 "2023-04-19T14:30:58Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![TheCedarPrince](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/thecedarprince/32/17323_2.png) [@TheCedarPrince](https://discourse.julialang.org/u/TheCedarPrince)\
**Post date:** [April 19, 2023, 2:30pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/1 "2023-04-19T14:30:58Z")

</div>

Hi folks,

I am giving a talk soon and something I want to highlight is the two language problem and how Julia addresses/solves this. Does anyone have a snippet of code that shows a simple Julia implementation of some function that can then be easily optimized (perhaps by types or other simple performance increases)? I would then use BenchmarkTools.jl to show the differences in timing – if that all makes sense.

Thanks folks!

~ tcp 🌳

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [April 19, 2023, 2:38pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/2 "2023-04-19T14:38:06Z")

</div>

I am not quite finding it, but I always liked @stevengj material on the matter.

The “Generating Vandermonde matrices” in [https://web.mit.edu/18.06/www/Fall17/1806/julia/Julia-intro.pdf](https://web.mit.edu/18.06/www/Fall17/1806/julia/Julia-intro.pdf) as an example.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [April 19, 2023, 2:47pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/3 "2023-04-19T14:47:00Z")

</div>

I was thinking of [julia-performance/1 - Is Julia fast?.ipynb at a1c77e92033c0ef3f58a360978ac2d3b08745ba8 · vchuravy/julia-performance · GitHub](https://github.com/vchuravy/julia-performance/blob/a1c77e92033c0ef3f58a360978ac2d3b08745ba8/lecture2%20-%20Performance%20Engineering/1%20-%20Is%20Julia%20fast%3F.ipynb) which I adapted from [18S096/Boxes-and-registers.ipynb at master · mitmath/18S096 · GitHub](https://github.com/mitmath/18S096/blob/master/lectures/lecture1/Boxes-and-registers.ipynb)

---

<div class="post-metadata">

**Author:** ![martin.auer](https://avatars.discourse-cdn.com/v4/letter/m/f19dbf/32.png) [@martin.auer](https://discourse.julialang.org/u/martin.auer)\
**Post date:** [April 19, 2023, 3:16pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/4 "2023-04-19T15:16:01Z")

</div>

One thing where Julia clicked for me in this respect was loop fusion/broadcasting [https://julialang.org/blog/2017/01/moredots/](https://julialang.org/blog/2017/01/moredots/)

---

<div class="post-metadata">

**Author:** ![s-broda](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/s-broda/32/3946_2.png) [@s-broda](https://discourse.julialang.org/u/s-broda)\
**Post date:** [April 19, 2023, 5:08pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/5 "2023-04-19T17:08:22Z")

</div>

Building on the `mysum` example, maybe it could be interesting to show how one can match the speed of the built-in by adding a simple decorator, and even surpass it by using LoopVectorization.jl, like in this [notebook](https://github.com/s-broda/presentation/blob/main/presentation.ipynb).

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [April 19, 2023, 5:25pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/6 "2023-04-19T17:25:10Z")

</div>

I would say that the best showcase is to take an existing example of a two language solution, e.g. a Matlab mex file, port it to Julia and show that you can get the same speed as the mex file and can more easily optimize it further. But this is rarely simple and very target specific as the audience should be familiar with the two language solution in question.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [April 19, 2023, 6:26pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/7 "2023-04-19T18:26:55Z")

</div>

by construction Two-Language problems aren’t likely to be simple, otherwise they’d just use the fast-lang in the two languages and come out fast in micro benchmarks

---

<div class="post-metadata">

**Author:** ![JonasWickman](https://avatars.discourse-cdn.com/v4/letter/j/9de0a6/32.png) [@JonasWickman](https://discourse.julialang.org/u/JonasWickman)\
**Post date:** [April 19, 2023, 6:31pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/8 "2023-04-19T18:31:55Z")

</div>

It’s really banal, but I would say that plotting capabilities are one of the prime candidates for why Julia is so good. Even if the problem is simple to implement in fast language like C or Fortran, in the end you will likely have to export a text file or something to read into Matlab/Python/R to see what happened.

---

<div class="post-metadata">

**Author:** ![tim.holy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tim.holy/32/52_2.png) [@tim.holy](https://discourse.julialang.org/u/tim.holy)\
**Post date:** [April 19, 2023, 6:40pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/9 "2023-04-19T18:40:18Z")

</div>

How about: “NumPy’s indexing rules are implemented in C, but Julia’s are written in Julia”? `base/abstractarrays.jl` becomes your complete example. (You could, e.g., count lines of code and compare to relevant portions of NumPy.)

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [April 19, 2023, 6:48pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/10 "2023-04-19T18:48:15Z")

</div>

It is not easy to find examples that impress every audience. For some people I show this kind of thing: computing some relative property of elements in a list, **but only for pairs in which one element satisfies one condition, and the other another condition**. This makes it at least not completely trivial to implement a vectorized version of the code, and tries to convey that the more complex or specific the calculation is, the more important is the standard performance of the straightforward implementations in a language.

For example, let us build a data set Pets, “Cats” or “Dogs”, in which the data structure carries the type of the Pet and its weight. (We can show off by putting units to the weights). We will compute the minimum difference between the weight of any dog or cat in this list:

```julia
using Unitful

struct Pet{T} 
    name::String
    weight::T
end

function min_diff_weight(pets)
    min_diff_weight = typemax(pets[1].weight)
    for i in 1:length(pets)-1
        if pets[i].name == "Dog"
            for j in i+1:length(pets)
                if pets[j].name == "Cat"
                    min_diff_weight = min(min_diff_weight, abs(pets[i].weight - pets[j].weight))
                end    
            end    
        end    
    end    
    return min_diff_weight
end

```

which takes for 10^4 pets:

```julia
julia> pets = [Pet(rand(("Dog","Cat")), rand()u"kg") for _ in 1:10^4];

julia> @btime min_diff_weight($pets)
  147.122 ms (0 allocations: 0 bytes)
3.578299201389967e-8 kg

```

We did nothing sophisticated there, we literally wrote the double loop that does exactly and explicitly what we wanted.

Now that in Python:

```python
import random

class Pet :
     def __init__ (self,name,weight) :
         self.name = name
         self.weight = weight

def min_diff_weight(pets) :
    min_diff_weight = float('inf')
    for i in range(0,len(pets)-1) :
        if pets[i].name == "Dog" :
           for j in range(i+1,len(data)) :
               if pets[j].name == "Cat" :
                   min_diff_weight = min(min_diff_weight, abs(pets[i].weight - pets[j].weight))
    return min_diff_weight

```

The codes are similar (although I would say Julia syntax is prettier…). But in Python this takes:

```python
In [10]: pets = [Pet(random.choice(("Dog","Cat")), random.random()) for _ in range(0,10_000)]

In [11]: %timeit min_diff_weight(pets)
3.07 s ± 14 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

```

Thus, it is about 20 times slower.

It is clear than anyone can come up with faster versions of these codes, using other tools, vectorizing, adding flags, numba, jits, etc. But with a basic knowledge of Julia you can get close to optimal performance for custom tasks like this without having to use any external tool or having to rethink your problem to adapt to those tools. When the problems become complex, this becomes more and more important.

---

<div class="post-metadata">

**Author:** ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)\
**Post date:** [April 19, 2023, 6:52pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/11 "2023-04-19T18:52:12Z")

</div>

> [@tim.holy](#):
>
> How about: “NumPy’s indexing rules are implemented in C, but Julia’s are written in Julia”?

Does this matter? It’s not like I can realistically change them either way.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [April 19, 2023, 7:11pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/12 "2023-04-19T19:11:23Z")

</div>

> [@JonasWickman](#):
>
> plotting capabilities

regrettably Julia plotting is way behind matplotlib. (not because matplotlib is good or anything, just because it’s been forever and people have covered every kind of plot, albeit lots of hacks like Artist object)

---

<div class="post-metadata">

**Author:** ![tim.holy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tim.holy/32/52_2.png) [@tim.holy](https://discourse.julialang.org/u/tim.holy)\
**Post date:** [April 19, 2023, 7:20pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/13 "2023-04-19T19:20:02Z")

</div>

I guess that’s up to you. But from the standpoint of lines of code, flexibility, generality (we have countless types of AbstractArrays, NumPy has one), yes, it does matter.

---

<div class="post-metadata">

**Author:** ![bertschi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bertschi/32/33462_2.png) [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Post date:** [April 19, 2023, 7:49pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/14 "2023-04-19T19:49:40Z")

</div>

Indeed, for basically the same reason I would consider Flux as a nice example … somehow it still strikes me as odd that [Tensorflow.jl](https://github.com/malmaud/TensorFlow.jl) is (was?) one of the largest Julia projects on Github in terms of LoC, i.e., just to expose the Tensorflow API you need 100k lines whereas the whole code of Flux fits into much less (had read it some years ago when it was still using Tracker and needed just 2.5k lines). Somehow Julia feels like cheating here 😉

---

<div class="post-metadata">

**Author:** ![simsurace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/simsurace/32/30216_2.png) [@simsurace](https://discourse.julialang.org/u/simsurace)\
**Post date:** [April 19, 2023, 8:18pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/15 "2023-04-19T20:18:32Z")

</div>

I think depending on the audience, the fact that you can write CUDA kernels in Julia via CUDA.jl is pretty nice.

---

<div class="post-metadata">

**Author:** ![Alec\_Loudenback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alec_loudenback/32/278_2.png) [@Alec\_Loudenback](https://discourse.julialang.org/u/Alec_Loudenback)\
**Post date:** [April 19, 2023, 9:33pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/16 "2023-04-19T21:33:32Z")

</div>

The team that made the famous black hole image has found a lot of benefit for a Julia implemenation of what previously had separate Python and C++ implemetnations:

> **[EnzymeCon ~ Day 2](https://www.youtube.com/live/NB7xUHQNox8?feature=share&t=18203)**

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [April 20, 2023, 9:57am UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/17 "2023-04-20T09:57:26Z")

</div>

I take inspiration from a discussion on an R data.table function underway [on Slack here](https://app.slack.com/client/T68168MUP/C674VR0HH/thread/C674VR0HH-1681522241.983199), to provide an example of how an enthusiast like me (but absolutely not an expert in programming) manages to implement (with some effort) some ideas to obtain in Julia the same result obtained with a function available in R. (In fact, I don’t know what R’s performance is on tables equivalent to the ones I tested. And I’m sure you can still improve what I got.)  
More precisely, the premise was this

> Is there an equivalent to R’s data.table `roll` on join in DataFrames.jl?The idea is, considering a left join, have the right table value cols being carry forward up to the next matching key …

Below is the temporal development of the ideas with tests for measuring the results

> **first attempt using DataFrames functions**
>
> ```julia
> using Dates, DataFrames, BenchmarkTools
> 
> tbl1 = [(;date) for date in Date(2022,1,1):Day(1):Date(2022,1,5)]
> tbl2 = [(date=Date(2022,1,2), var1=rand()), (date=Date(2022,1,5), var1=rand())]
> tbl3 = [(date=Date(2022,1,1), var2=rand()), (date=Date(2022,1,3), var2=rand())]
> DF1=DataFrame(tbl1)
> DF2=DataFrame(tbl2)
> DF3=DataFrame(tbl3)
> df123=outerjoin(DF1,DF2, DF3, on=:date, order=:left)
> filldown(x)=accumulate((p,f)->coalesce(f,p), x)
> @btime transform(df123, [:var1, :var2].=>filldown.=>[:var1, :var2])
> # 27.200 μs (258 allocations: 14.50 KiB)
> # 5×3 DataFrame
> 
> DF1=deleteat!(DataFrame(tbl1),[2])
> 
> @btime leftjoin(DF1,transform(sort(outerjoin(DF1,DF2, on=:date),:date),:var1=>filldown=>:var1), on=:date) 
> 
> ```

> **a solution using the eclectic FlexiJoins package**
>
> ```julia
> 
> tbl1 = [(;date) for date in Date(2000,1,1):Day(1):Date(2022,1,5)];
> tbl2 = (date=[d for d in Date(2000,1,1):Day(1):Date(2022,1,5)], var=rand(8041));
> df1=DataFrame(deleteat!(tbl1, sort(unique(rand(1:7041,1000)))));
> df2=deleteat!(DataFrame(tbl2), sort(unique(rand(1:7041,4000))));
> t2=copy.(eachrow(df2));
> 
> using FlexiJoins
> outFJ=last.(FlexiJoins.innerjoin((A=tbl1, B=t2), by_pred(:date, ≥, :date), multi=(B=closest,)).B)
> 
> ```

> **a solution using only basic functions**
>
> ```julia
> roll_ssl(c1,c2)=[searchsortedlast((c2), k) for k in c1]
> 
> @btime begin
> issl=roll_ssl(df1[!,1],df2[!,1])
> df1.var=similar(df1.date,Float64)
> si=findfirst(==(1),issl)
> df1.var[1:si-1] .= NaN
> df1.var[si:end]=df2.var[issl[si:end]]
> end
> # 243.500 μs (16 allocations: 222.64 KiB)
> # 7109-element Vector{Float64}:
> 
> ```

finally a solution using only for/while loop, which, with some effort, allows to obtain an improvement of at least 10 times

> **roll**
>
> ```julia
> 
> function roll(c1,c2)
> j=1
> g=similar(c1,Int)
> li=1
> while(c1[li]<c2[1])
> g[li]=0
> li+=1
> end
> for i in li:length(c1)
> if (j == length(c2))
> g[i:end].=j
> break
> else
> if c1[i] < c2[j+1] 
> g[i]=j
> else
> while (j < length(c2)) && (c2[j+1]<=c1[i])
> j+=1
> end
> g[i]=j
> end
> end
> end
> g
> end
> 
> @btime begin
> iw=roll(df1.date,df2.date)
> df1.var=similar(df1.date,Float64)
> si=findfirst(==(1),iw)
> df1.var[1:si-1] .= NaN
> df1.var[si:end]=df2.var[iw[si:end]]
> end
> # 26.300 μs (16 allocations: 222.64 KiB)
> # 7109-element Vector{Float64}:
> 
> ```

---

<div class="post-metadata">

**Author:** ![StatisticalMouse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/statisticalmouse/32/43370_2.png) [@StatisticalMouse](https://discourse.julialang.org/u/StatisticalMouse)\
**Post date:** [April 20, 2023, 5:32pm UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/18 "2023-04-20T17:32:47Z")

</div>

It would be a lot easier to show instances where the two-language problem is a problem, like with python and Spark or Flink, but this doesn’t really answer the question as I don’t see Julia as a drop-in replacement for these tools.

Julia could of course fit these domains with more work, but Rust-based solutions are further, and they have plenty of momentum.

---

<div class="post-metadata">

**Author:** ![sylvaticus](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sylvaticus/32/203883_2.png) [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Post date:** [April 21, 2023, 5:59am UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/19 "2023-04-21T05:59:45Z")

</div>

As others have said, the solved “two language oroblem” are by definition untrivisl cases.

What you can do is show screenshots of the github pages with the charts for Numpy, pandas, scikit-leatn of the language they are written in… and then show the same beautiful “100% Julia” charts for DataFrames, CSV, Flux …

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [April 21, 2023, 6:55am UTC](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659/20 "2023-04-21T06:55:07Z")

</div>

Enzyme is:  
 ![image](https://global.discourse-cdn.com/julialang/original/3X/a/9/a987f90cb19f95e34b03c7214fe5530cbd0b14b9.png)

Unlike (what eventually would be?) Rust Enzyme, Julia is purely a front end of Enzyme, so it’s really a two-language solution in that regard.

Like, in this particular case, if C++ adopts Enzyme, it would still be 1 language solution and presumably performance will be better.

[Next page](https://discourse.julialang.org/t/what-is-a-simple-example-of-julia-solving-the-two-language-problem/97659.md?page=2)
