# Can Julia really be used as a scripting language? (Performance)

**URL:** <https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384>\
**Category:** Performance\
**Created:** [May 29, 2020, 1:06am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384 "2020-05-29T01:06:51Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 29, 2020, 1:06am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/1 "2020-05-29T01:06:51Z")

</div>

I’ve been reading the docs, and getting pretty excited about Julia’s potential as a replacement for Ruby in my workflows, however I very quickly seem to have run into a brick wall once trying to actually use it: it’s incredibly slow.

This is sort of surprising to me as I’m using only versions greater than 1.0 and have now switched to 1.4.1 (latest on downloads page). At first the REPL loading was a problem, but 1.4.1 fixes that. My current test case is a two-line script `import CSV` and then the for loop over `CSV.File` and doing nothing in the loop. I have a 4-line CSV file I’m testing with. The script takes almost 30 seconds to run on my desktop computer.

Is this just a case of my use case not being aligned with the purpose of Julia? I know the history here is for long-running “data science” type operations that take days, etc. The REPL speed increase from ~1.0 to 1.4.1 gives me hope that the scripting use case is in fact interesting to someone other than me, but I’m basically wondering how much. Like, could I start using Julia today and hope than in a few years it’ll be usable, or is it just such antithesis to the main purposes/uses of the environment that I should be looking elsewhere?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 29, 2020, 1:12am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/2 "2020-05-29T01:12:32Z")

</div>

> [@singpolyma](#):
>
> Is this just a case of my use case not being aligned with the purpose of Julia?

It’s well known that Julia needs to compile the functions at the start. For scripting, the best that can be done is to use the lowest optimization level `julia -O 0` and see if that helps. But I do know ppl who use Julia for scripting, so they may have other ways to improve the exp as well.

The other way is to compile binary using PackageCompilers.jl but that may not be possible.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [May 29, 2020, 1:22am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/3 "2020-05-29T01:22:12Z")

</div>

It might be wroth trying to precompile your packages.  
In the REPL, type `]precompile`, which will take a minute or two, but might speed things up afterwards.

---

<div class="post-metadata">

**Author:** ![PetrKryslUCSD](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/petrkryslucsd/32/215825_2.png) [@PetrKryslUCSD](https://discourse.julialang.org/u/PetrKryslUCSD)\
**Post date:** [May 29, 2020, 2:34am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/4 "2020-05-29T02:34:38Z")

</div>

Well, here is some data:

```julia
pkonl@TeenyTiny MINGW64 ~/Documents
$ time ~/AppData/Local/Programs/Julia/Julia-1.5.0-DEV/bin/julia.exe -O0 do.jl
a=missing, b=1, c=1.0
a=missing, b=2, c=2.0
a=missing, b=3, c=3.0

real 0m13.356s
user 0m0.015s
sys 0m0.015s

pkonl@TeenyTiny MINGW64 ~/Documents
$ cat do.jl data.csv
using CSV

for row in CSV.File("data.csv")
    println("a=$(row.a), b=$(row.b), c=$(row.c)")
end
a,b,c,col4,col5,col6,col7,col8
,1,1.0,1,one,2019-01-01,2019-01-01T00:00:00,true
,2,2.0,2,two,2019-01-02,2019-01-02T00:00:00,false
,3,3.0,3.14,three,2019-01-03,2019-01-03T00:00:00,true

```

With Julia 1.5, precompiled.

---

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 29, 2020, 2:49am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/5 "2020-05-29T02:49:01Z")

</div>

Thanks for that. So it’s not just me or my settings, even doing everything “right” on the newest version it’s very slow. Which is fine for many use cases of course, I’m not trying to be judgemental here, I’m just curious if this is considered an issue that should be worked on eventually, or just “the cost of the way we do things”.

Or maybe the CSV library is just very big and this isn’t often an issue? Hello world isn’t too slow (about the same speed as ruby, faster than runhaskell in my tests)

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [May 29, 2020, 3:16am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/6 "2020-05-29T03:16:06Z")

</div>

If you do not anything fancy but just to read CSVs then you can use [`DelimitedFiles.readdlm`](https://docs.julialang.org/en/v1/stdlib/DelimitedFiles/#DelimitedFiles.readdlm-Tuple%7BAny,AbstractChar,Type,AbstractChar%7D) that exists in the standard library (i.e., you do not need to install any package, just `import` it). Yes, there were other people using `CSV.jl` that had the same problem with startup problems recently.

Also, I run my Julia scripts with options `-O0 --compile=min`, you can even use `--compile=no`. The flag `--help-hidden` also shows a `--trace-compile` to help you find why it takes so much time to start. However, even with all of this, “first time to plot” is a known issue that is being fought by Julia developer team. My workaround for exploratory data analysis is using a Jupyter notebook with package `Revise` imported and letting it open all the time, so after the first calls to each function, the code runs blazing fast.

---

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 29, 2020, 3:43am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/7 "2020-05-29T03:43:35Z")

</div>

> [@Henrique\_Becker](#):
>
> `--compile=min`

Holy shit, thank you! This cuts my time down from 30s to 2.5s with just this switch. It’s not great, but it’s viable for sure.

* * *

Switching from `CSV` to `DelimitedFiles` is probably ok for my case, and gets me to sub-second so seems worth it for now. Thanks for the tip.

---

<div class="post-metadata">

**Author:** ![leethargo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leethargo/32/6004_2.png) [@leethargo](https://discourse.julialang.org/u/leethargo)\
**Post date:** [May 29, 2020, 7:15am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/8 "2020-05-29T07:15:00Z")

</div>

For a related use case (call a simple Julia script repeatedly), where the start-up and compilation time would dominate the actual runtime, I’ve thought about using an implicit client/server approach to reuse a process.

This is inspired by [Emacs Server](https://www.gnu.org/software/emacs/manual/html_node/emacs/Emacs-Server.html#Emacs-Server), which allows to run `emacsclient <file>` in a shell, but actually open the file in the long-running Emacs instance. This works by starting an Emacs Server on start-up of Emacs, then using the `emacsclient` instead of `emacs` for later calls.

I imagine the following (simpler) workflow with Julia: There is a bash script `juliaserver` that can be used with `juliaserver script.jl`. The first time it is called, it will open a separate background process and use that to execute the file (using `include`). The next time, the existing process will be detected (e.g. by locating a specific file) and a short message with the call arguments is sent there, to be executed.

There are some questions, still, such as: will the execution of later calls be influenced by previous calls, because of remaining variables in scope, or invalid redefinition of types? How long should the process be kept running idle? What environment should be active? So, it’s not quite obvious how to decide on the details.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [May 29, 2020, 8:01am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/9 "2020-05-29T08:01:41Z")

</div>

> [@singpolyma](#):
>
> I’m just curious if this is considered an issue that should be worked on eventually, or just “the cost of the way we do things”.

I think it’s both. It is a known cost of how Julia works, with its jit compiler and aggressive type specialization. But it is also a _very_ actively talked about issue, that has a great part of the core devs’ attention. I believe that compiler performance is close to the top of their list of priorities.

---

<div class="post-metadata">

**Author:** ![lwabeke](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lwabeke/32/4005_2.png) [@lwabeke](https://discourse.julialang.org/u/lwabeke)\
**Post date:** [May 29, 2020, 9:04am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/10 "2020-05-29T09:04:44Z")

</div>

Recently @kristoffer.carlsson gave a nice webinar on PackageCompiler. I’m not sure if it was recorded and is still available.  
But that showed that how the startup issues can be reduced significantly.  
If you know you need to use CSV.jl in a scripting environment, it might be worthwhile to add that to your julia sysimage. And launching that sysimage as your scripting environment. That should get the launching of Julia and the using CSV steps to be sub-second.

The downside is some setup work and more manual version stepping of the packages, but for your scripting environment, you probably want stable versions and don’t want things to change every few days.

Further speedup can be obtained by compiling the specific functions needed, but depending on the complexity of the code that gets JIT compiled, that might not even be necessary.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [May 29, 2020, 1:55pm UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/11 "2020-05-29T13:55:09Z")

</div>

You know you are basically describing a CLI for Jupyter notebooks, no?

---

<div class="post-metadata">

**Author:** ![leethargo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leethargo/32/6004_2.png) [@leethargo](https://discourse.julialang.org/u/leethargo)\
**Post date:** [May 29, 2020, 2:26pm UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/12 "2020-05-29T14:26:22Z")

</div>

That’s a good point. Maybe the Jupyter infrastructure can be reused for this, rather than building something from scratch with `Socket`s etc.

But it should be more implicit, with the _kernel_ started automatically, and each new call using a fresh _session_, if possible. So we would only want to avoid recompilation, not keep any state between calls.

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [May 29, 2020, 2:37pm UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/13 "2020-05-29T14:37:53Z")

</div>

Well, the kernel being started automatically can be done by a wrapping script that checks if it is running and if it is not, then starts it. The call does not need to be in a fresh session, it just needs to be wrapped in a function or some other scope to not leak variables unless you are really worried about some kind of global state.

---

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 29, 2020, 3:01pm UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/14 "2020-05-29T15:01:03Z")

</div>

I wonder if more of the compilation could be cached. If a precompiled  
sysimage could speed it up, could the parts for that for each module not  
be created and cached automatically? I believe Guile does something like  
this.

In any case, I think the compile=min, while not fast, is fast enough for  
now especially if there’s hope for it to get even better with time.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [May 30, 2020, 2:22am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/15 "2020-05-30T02:22:14Z")

</div>

> [@singpolyma](#):
>
> is fast enough for  
> now especially if there’s hope for it to get even better with time.

On the better with time point, [see here](https://discourse.julialang.org/t/compiler-work-priorities/17623). If you want to read lots more about this, search “time to first plot,” mentioned above, which is a shorthand for this issue (plotting packages are notorious for long compilation times).

But I would also encourage you to explore other workflows. I come from bioinformatics, where everything is a script, so I totally get the inertia. But now I use the REPL and Atom, for most development, and only write scripts for long running precesses where compile time is a tiny fraction. Think about it this way: when you’re coding interactively, use the interactive tools.

A couple workflows have been mentioned here, there are also great ways to use Atom or VS code as development environments. I usually start up Julia first thing, run a script that loads my packages (esp. Revise.jl), get a cup of coffee, and then don’t worry about compilation for the rest of the day.

---

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 30, 2020, 2:49am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/16 "2020-05-30T02:49:19Z")

</div>

> [@kevbonham](#):
>
> I usually start up Julia first thing, run a script that loads my packages (esp. Revise.jl), get a cup of coffee, and then don’t worry about compilation for the rest of the day.

I totally understand that for the “data science” use case Julia is advertised for, this is reasonable. If I want to ship a script to users, though, it needs to work as a script 🙂

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [May 30, 2020, 7:25am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/17 "2020-05-30T07:25:53Z")

</div>

Perhaps automatically wrap each script in a module?

---

<div class="post-metadata">

**Author:** ![singpolyma](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singpolyma/32/14898_2.png) [@singpolyma](https://discourse.julialang.org/u/singpolyma)\
**Post date:** [May 30, 2020, 11:06am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/18 "2020-05-30T11:06:39Z")

</div>

I can wrap my code in a module. What would that accomplish? Most of the  
code is in dependencies which are modules already.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [May 30, 2020, 11:15am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/19 "2020-05-30T11:15:55Z")

</div>

I was suggesting that as a solution to give each script their own namespace if run on a hypothetical long running julia-server.

That serve approach would have the advantage of only needing to compile the dependencies once. As long as it keeps running as a background process, subsequent runs of the script won’t have to recompile.

That won’t help you if the folks you’re sending it to only run it once.

---

<div class="post-metadata">

**Author:** ![Per](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/per/32/10387_2.png) [@Per](https://discourse.julialang.org/u/Per)\
**Post date:** [May 30, 2020, 11:57am UTC](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384/20 "2020-05-30T11:57:07Z")

</div>

Maybe somebody could distribute a “batteries included” Julia binary that uses [PackageCompiler](https://github.com/JuliaLang/PackageCompiler.jl) to allow a number of the most popular packages to start up instantly.

(This would be aimed at people who simply want to run scripts that they get sent to them, and who are not bothered by not having the very latest versions of packages.)

[Next page](https://discourse.julialang.org/t/can-julia-really-be-used-as-a-scripting-language-performance/40384.md?page=2)
