# Why is Python, not Julia, still used for most state-of-the-art AI research?

**URL:** <https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896>\
**Category:** Offtopic\
**Tags:** knet, flux, machine-learning, mlj, sciml\
**Created:** [September 1, 2020, 1:49pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896 "2020-09-01T13:49:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 1:49pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/1 "2020-09-01T13:49:21Z")

</div>

I know e.g. ALBERTA, a variant of BERT, has been recreated in Julia, and AlphaZero, but all the (recent) papers I read, I see are still based on Python code, if the language is known to me (e.g. GPT-2 and updated GPT-3). [SciML, merging neural networks and PDEs, maybe the the exception.] What (other) counterexamples are there?

Ok, often I’m not sure what language is used, e.g. for interesting paper from July:

> **[Self-supervised Neural Architecture Search](https://arxiv.org/abs/2007.01500)**
>
> Neural Architecture Search (NAS) has been used recently to achieve improved performance in various tasks and most prominently in image classification. Yet, current search strategies rely on large labeled datasets, which limit their usage in the case...

and: Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense  
[https://arxiv.org/pdf/2004.09044.pdf](https://arxiv.org/pdf/2004.09044.pdf)

My hope was that, as we need, and are making newer algorithms all the time, obsoleting older ones, people would use Julia, and Python would simply get outdated that way.

Good blog:

> **[Artificial intelligence: The dark matter of computer vision - TechTalks](https://bdtechtalks.com/2020/06/01/artificial-intelligence-computer-vision-fpicu/)**
>
> AI researchers at MIT and UCLA discuss the need to address the "dark matter" of computer vision, the things that aren't visible in pixels.

> _This article is part of our [reviews of AI research papers](https://bdtechtalks.com/tag/ai-research-papers/), a series of posts that explore the latest findings in artificial intelligence._

> **[An AI-written blog highlights bad human judgment on GPT-3 - TechTalks](https://bdtechtalks.com/2020/08/24/ai-blog-gpt-3-fake-news/)**
>
> A blog written by GPT-3 triggered a lot of hype in the media. But most stories paint an incorrect picture of advances in AI.

> the AI-generated blog’s popularity on tech-focused Hacker News forum as a measure that GPT-3 had managed to fool readers that a human had written the posts.

 ![Hacker News post on AI-generated blog written by GPT-3](https://global.discourse-cdn.com/julialang/original/3X/5/f/5f0ec2c0bb9e3245b0fdf327ba714810fb9d8a5e.png)

> A post written by GPT-3 made it to the top of the Hacker News forum
> 
> The post has indeed received 198 points and 71 comments. “What most commenters didn’t realize: The post was generated entirely by artificial intelligence,” [Business Insider](https://www.businessinsider.com/fake-ai-generated-gpt3-blog-hacker-news-2020-8?r=DE&IR=T) wrote.

EDIT: I did find from the main guy behind Knet.jl (but the paper doesn’t mention if Julia or other language used):

> **[Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional...](http://www.denizyuret.com/2020/03/bilingunet-image-segmentation-by.html)**
>
> İlker Kesen, Ozan Arkan Can, Erkut Erdem, Aykut Erdem, Deniz Yuret. June 20, 2022. Best paper at the 5th Multimodal Learning and Application...

> **Abstract:** We present BiLingUNet, a state-of-the-art model for image segmentation using referring expressions. BiLingUNet uses language to customize visual filters and outperforms approaches that concatenate a linguistic representation to the visual input. […]

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [September 1, 2020, 1:58pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/2 "2020-09-01T13:58:03Z")

</div>

The content of the PyTorch repo sheds some light on this:

![Screen Shot 2020-09-01 at 9.56.56 AM](https://global.discourse-cdn.com/julialang/original/3X/b/a/ba6de217ca90900d7f301fe2329be79e8334a578.png)

---

<div class="post-metadata">

**Author:** ![anon37204545](https://avatars.discourse-cdn.com/v4/letter/a/439d5e/32.png) [@anon37204545](https://discourse.julialang.org/u/anon37204545)\
**Post date:** [September 1, 2020, 1:59pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/3 "2020-09-01T13:59:08Z")

</div>

From a thread about CSV:

> [@CSV Reader Benchmarks: Julia Reads CSVs 10-20x Faster than Python and R](https://discourse.julialang.org/t/csv-reader-benchmarks-julia-reads-csvs-10-20x-faster-than-python-and-r/41974/54):
>
> Other than a few missing features, there are projects in both languages that are already there: [python](https://github.com/SciML/diffeqpy), [R](https://github.com/SciML/diffeqr).

> [@CSV Reader Benchmarks: Julia Reads CSVs 10-20x Faster than Python and R](https://discourse.julialang.org/t/csv-reader-benchmarks-julia-reads-csvs-10-20x-faster-than-python-and-r/41974/61):
>
> My recollection of history there was Chris was trying to publish a paper about DifferentialEquations.jl. And it got rejected because “it is in an obsure language noone will use it”, and in a weekend of pique, Chris coded up the python and R wrappers. So that one reviewer probably did more to advance the cause of DE solving in R and Python than many R or Python programmer working in the area ever have.

It seems it’s common that in order to get a paper more valued, you have to use the current standard language of the field (i.e. Python).

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 2:01pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/4 "2020-09-01T14:01:22Z")

</div>

You imply it’s really C++, fair enough, my main point was, why not Julia? And maybe you imply that combination is good enough, Julia not needed. I think most users use Python, without knowing or caring C++ powers it, not just in AI. I suppose (some of the) research people really do stuff in C++ though.

Also why not then use [GitHub - boathit/JuliaTorch: Using PyTorch in Julia Language](https://github.com/boathit/JuliaTorch) wrapper, rather than Python directly and C++ indirectly?

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [September 1, 2020, 2:52pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/5 "2020-09-01T14:52:19Z")

</div>

I don’t see why this is a mystery. Everybody uses python because everybody uses python. It’s what they’re used to, and everyone they collaborate with uses it.

There’s also lots of sunken cost, lots of available tools, and changing language incurs a cost.

That’s how inertia works. It’s nothing to do with python or Julia, as languages.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 2:54pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/6 "2020-09-01T14:54:48Z")

</div>

E.g. binary networks, 1-bit weights and activations, seems like would have been easier to develop in Julia (than in Lua, so people to migrate from one language to newer), understandable they didn’t in 2016: [https://arxiv.org/pdf/1602.02830.pdf](https://arxiv.org/pdf/1602.02830.pdf)

> Torch7 and Theano framework […]  
> Last but not least, we wrote a binary matrix multiplication GPU kernel with which it is possible to run our MNIST BNN 7 times faster than with an unoptimized GPU kernel, without suffering any loss in classification accuracy.

In Python/PyTorch as late as 16 days ago:

> **[GitHub - liuzechun/ReActNet: ReActNet: Towards Precise Binary NeuralNetwork...](https://github.com/liuzechun/ReActNet)**
>
> ReActNet: Towards Precise Binary NeuralNetwork with Generalized Activation Functions. In ECCV 2020. - GitHub - liuzechun/ReActNet: ReActNet: Towards Precise Binary NeuralNetwork with Generalized Ac...

> This is the pytorch implementation of our paper [“ReActNet: Towards Precise Binary NeuralNetwork with Generalized Activation Functions”](https://arxiv.org/abs/2003.03488), published in ECCV 2020.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [September 1, 2020, 3:02pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/7 "2020-09-01T15:02:42Z")

</div>

> [@DNF](#):
>
> That’s how inertia works. It’s nothing to do with python or Julia, as languages.

I couldn’t agree more on this. People don’t care about the language, they just use whatever everybody around them is using already.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 3:12pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/8 "2020-09-01T15:12:33Z")

</div>

I do see people in the Julia community doing lots of research, mostly non-AI (and yes, SciML), and I know about the ML/AI Julia packages, the infrastructure. I suppose they could be used by now instead of e.g. Python/C++, and maybe people are using (just not everyone aware of, at least I’ve not seen any major AI original research done using Julia). I was thinking, do they currently have some limitations that make Python justified, or is it simply ignorance of the Julia alternatices and/or inertia?

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [September 1, 2020, 3:28pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/9 "2020-09-01T15:28:32Z")

</div>

I think the non-technical side of things has been summarized already, so let me add a couple of technical issues that Python frameworks still have the edge in:

1. Distributed and multi-gpu training: this has unfortunately become a must for certain streams of research in CV, NLP and deep RL. CliMA is the only public project I know of that has distributed GPU support, but it’s not DL and the only way to do something similar for DL right now is to implement your own framework from scratch on top of MPI (or resuscitate NCCL.jl, but that’s even less likely).
2. API coverage: for example, over half of my lab is MATLAB refugees who really only use Python for DL model training. They are not in a position to write custom kernels and rely on the provided  
kitchen sink to hack together new layers. You can see this browsing through the repos of many research papers: the publication promises a theoretical speedup over, say, quadratic self attention, but the actual implementation is an inefficient tangle of whatever `torch.tensor` ops were necessary to get things running.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 3:51pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/10 "2020-09-01T15:51:31Z")

</div>

> [@ToucheSir](#):
>
> Distributed and multi-gpu training: this has unfortunately become a must for certain streams of research in CV, NLP and deep RL. CliMA is the only public project I know of that has distributed GPU support, but it’s not DL and the only way to do something similar for DL right now is to implement your own framework from scratch on top of MPI

Hmm, yes, multi-GPU was a blind spot to me, I see it done as far back as 2016, but not specific to ANNs/DL:

> **[Multiple-GPU Parallelism on the HPC with Julia - Stochastic Lifestyle](https://www.stochasticlifestyle.com/multiple-gpu-on-the-hpc-with-julia/)**
>
> This is the exciting Part 3 to using Julia on an HPC. First I got you started with using Julia on multiple nodes. Second, I showed you how to get the code running on the GPU. That gets you pretty far. However, if you got a trial allocation on...

> [@How to use multiple GPUs correctly?](https://discourse.julialang.org/t/how-to-use-multiple-gpus-correctly/29986):
>
> Hi all, I am trying to change my GPU code to use multiple GPUs due to the memory limit of a single GPU. I found an example ([Multiple-GPU Parallelism on the HPC with Julia | juliabloggers.com](https://www.juliabloggers.com/multiple-gpu-parallelism-on-the-hpc-with-julia/)) and the basic idea is split the whole data into different parts. store each part of the data using CuArrays in different GPU cards. launch the kernels asynchronously in different GPUs with local data. So I have done a test with the following code using CuArrays, CUDAnative using BenchmarkTools N = 2 …

> [@ToucheSir](#):
>
> or resuscitate NCCL.jl

Is it for sure dead? This one or at JuliaGPU (both updated recently)? [GitHub - vchuravy/NCCL.jl: A Julia wrapper for the NVIDIA Collective Communications Library.](https://github.com/vchuravy/NCCL.jl)

There’s interesting work being done to scale NNs down not just up (as with GPT-3), both for NLP and computer vision. Still, GPT-3 is huge (ALBERTA I mentioned much smaller), so multi-GPU seems needed for sure (at least for good NLP now).

I’m curious, if the network itself doesn’t need to be that big (say fits on one memory), but the problem is the dataset/training, what happens if you spit it 2 or N ways and train independently, can you in general (or say for images only) combine two such trained networks? Isn’t that what people call minibatching? I could see it maybe not working for NLP.

And can you simply use:

> **[GitHub - horovod/horovod: Distributed training framework for TensorFlow,...](https://github.com/horovod/horovod)**
>
> Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet. - GitHub - horovod/horovod: Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.

> Horovod is a distributed deep learning training framework for TensorFlow, Keras, PyTorch, and Apache MXNet

E.g. Julia has by now (for a long time, while it’s not the post popular framework, for Julia or otherwise) official support for MXNet, and as I posted, there’s a PyTourch wrapper, while the Tensorflow one is a bit outdated.

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [September 1, 2020, 4:00pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/11 "2020-09-01T16:00:11Z")

</div>

> [@giordano](#):
>
> I couldn’t agree more on this. People don’t care about the language, they just use whatever everybody around them is using already.

There’s a bit more to it. Currently, out of a dozen or more colleagues, not one has any understanding whatsoever of what makes Julia compelling. If they have guess, it’s that Julia allows writing optimized numerical code, maybe some kind of modern C++. As Viral and others have pointed out, it seems very difficult to communicate the advantages laid out in this thread, say in slides or a blog post. Actually, I am thinking about how best to do a presentation about it now.  
Having said that, there still may be reasons to stick with python. For example, your end users only use computers because they have to and have no interest in programming.  
So inertia may explain it. But, for the most part, people have no idea what they are missing.  
I do find that people who have experience in more expressive statically compiled languages are painfully aware of some of the limitations of python, which at least allows you to imagine that Julia might have something to offer.

---

<div class="post-metadata">

**Author:** ![giordano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/giordano/32/2166_2.png) [@giordano](https://discourse.julialang.org/u/giordano)\
**Post date:** [September 1, 2020, 4:03pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/12 "2020-09-01T16:03:35Z")

</div>

> [@jlapeyre](#):
>
> But, for the most part, people have no idea what they are missing.

Yes, because they don’t care about the language they use, be it R, Python or Julia.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 4:06pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/13 "2020-09-01T16:06:26Z")

</div>

> [@jlapeyre](#):
>
> there still may be reasons to stick with python. For example, your end users only use computers because they have to and have no interest in programming.

Yes, but I doubt that for “state-of-the-art AI research”. I could be wrong. If you meant more generally, or outside AI, say web programming, then sure. Then I’m not sure Julia has a killer advantage. But for AI, and I really meant ML, then I thought Julia had such an advantage based on (I realize many ML people are doing not too advanced stuff, but I’m not thinking of those): [Building a Language and Compiler for Machine Learning](https://julialang.org/blog/2018/12/ml-language-compiler/)

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [September 1, 2020, 4:11pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/14 "2020-09-01T16:11:33Z")

</div>

> [@Palli](#):
>
> Yes, but I doubt that for “state-of-the-art AI research”.

Yes, you are correct. I was inadvertently broadening the topic. But, I’m not thinking of web frameworks, but rather other scientific computing. (But, I guess people who write web framework stuff in Julia, would have a similar story)

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [September 1, 2020, 4:12pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/15 "2020-09-01T16:12:16Z")

</div>

Agreed. On top of this, evangelism is both difficult and expensive. The major Python frameworks have corporations with extensive name recognition and oodles of cash backing entire conferences, YouTube channels with highly-produced videos, blogs and more. Even big corporate backing is not a guarantee of breakout success: Swift for TensorFlow remains somewhat obscure despite the initial buzz.

With all that said, let me plug [https://github.com/JuliaCommunity/ML-Coordination-Tracker](https://github.com/JuliaCommunity/ML-Coordination-Tracker), where we’re trying to identify and fill in gaps in the ecosystem. If you’re interested in a more grassroots approach for increasing adoption, feel free to open an issue or hop on Zulip 🙂

---

<div class="post-metadata">

**Author:** ![chakravala](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chakravala/32/6832_2.png) [@chakravala](https://discourse.julialang.org/u/chakravala)\
**Post date:** [September 1, 2020, 4:15pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/16 "2020-09-01T16:15:42Z")

</div>

In my field of geometric algebra, my Julia package has 100 more stars than the equivalent python package. However, pretty much _all_ the open source collaborative effort is going the python package, and zero of the collaboration into the Julia package. So even though my Julia package is more (apparently) popular at this point, the python package already has an established community of contributors (a group of them, we all stay in contact on our communication server and share all our work accross languages). However, I don’t actually mind that the rest of my geometric algebra developer friends are focusing on the python language instead of Julia. We formed a unified geometric algebra community, which is independent of choice of language. There are a bunch of developers doing it in other languages also, but for this sake I focus on python vs Julia.

So, I think we can all get along, no matter what language they are using. In the geometric algebra community, I am happy there is such a diverse choice of languages, including Julia and Python. It’s great that there are people using python independently, because it’s a separate effort I can keep learning from.

I don’t try to convince my algebra friends to switch to Julia anymore, what language they use is up to them and their situation. Yea, I will recommend Julia, but I don’t mind if they choose to neglect the Julia language.

The python project is probably easier to contribute to for them because it doesn’t have obscure usages of metaprogramming features like in Julia, so it gives an easier entry point for a complicated project.

I am really good friends with the developers of the python variant of my geometric algebra package, and talk to them by phone fairly often too (about science, not programming, language agnostic). Our shared lanuage is mathematics.

Anyways, my point is that scientific collaboration should not put so much emphasis in the language choice. Instead, the most fruitful discussions happen if you create unified communities beyond only Julia. Let other people find their way to Julia on their own, whenever they are ready.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [September 1, 2020, 4:33pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/17 "2020-09-01T16:33:03Z")

</div>

Your comment and my response seems off-topic (still interesting, and maybe relevant in general about collaboration), regarding AI, but I really like to know if, with me just seeing Hyperbolic Deep Learning idea/thread today.

> [@chakravala](#):
>
> In my field of geometric algebr

I’m aware of you and your packages, and I must admit, they’re above my paygrade. I understand physics at a level I’m ok with, but suppose I need to look into e.g. Clifford (Grassmann, Lie?) algebra, if I really want to understand the physics I want to know more about (e.g. spinnors, twistors and Geometric Unity).

---

<div class="post-metadata">

**Author:** ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Post date:** [September 1, 2020, 4:33pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/18 "2020-09-01T16:33:55Z")

</div>

There are countries that still do not use the metric system, you figure it out.

---

<div class="post-metadata">

**Author:** ![dlakelan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dlakelan/32/8491_2.png) [@dlakelan](https://discourse.julialang.org/u/dlakelan)\
**Post date:** [September 1, 2020, 4:40pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/19 "2020-09-01T16:40:48Z")

</div>

CMU Common Lisp was available for x86 back in the mid 90’s It had a LOT of the advantages that Julia has: a good compiler, macros, multi-methods in CLOS, etc. Then the SBCL variant of CMUCL came along and it got more maintainable and improved further.

I mean, no question Julia has many advantages such as a syntax that people like, and probably a fancier compiler that can specialize things based on types without type annotations, but in 1999 you could write code that would blow Python out of the water, yet it never caught on in popular usage.

I think the answer is high end languages like Common Lisp and Julia are really _for people with a high end view of computing_ and there are just far fewer of them than there are general people who do some coding.

What’s the advantage of macros if you don’t really understand what they mean and how they work? What’s the advantage of a compiler that can do all kinds of type inference so duck-typing is a thing, if you don’t understand why that’s a thing and you do concrete type annotations so everything has to be a Float64? etc etc.

A lot of what makes Julia so amazing is that it attracts amazing people who understand how Julia code could be made amazing.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [September 1, 2020, 4:47pm UTC](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896/20 "2020-09-01T16:47:08Z")

</div>

> [@Palli](#):
>
> Yes, but I doubt that for “state-of-the-art AI research”. I could be wrong. If you meant more generally, or outside AI, say web programming, then sure. Then I’m not sure Julia has a killer advantage. But for AI, and I really meant ML, then I thought Julia had such an advantage based on (I realize many ML people are doing not too advanced stuff, but I’m not thinking of those): [Building a Language and Compiler for Machine Learning](https://julialang.org/blog/2018/12/ml-language-compiler/)

This is a compelling argument, but I don’t think it holds in practice (yet) outside of niches like SciML. For better or worse, many ML research papers just don’t care about training time, instead choosing to focus on (theoretical) advantages or evaluation metrics while throwing more compute/data at the problem (case in point, anything that uses a TPU cluster). The minority of publications that _do_ care either  
a) overlap with SciML and neural diffeqs already,  
b) are concerned with (device-side) inference (which Flux/Knet don’t address at all), or  
c) are engineering-focused and written by authors with plenty of C++/CUDA experience.

Now, let me play devil’s advocate and address that blog post specifically. Here’s a question: how many of those points does JAX/XLA not handle already? The only one I can think of is custom kernels, but again the trend of current ML research is not to implement custom kernels. On the other hand, Flux’s TPU support is now completely out of date and automatic batching practically does not exist. Likewise, broadcast fusion is great but ML-specific optimizers like XLA can go beyond that and pull off even more aggressive optimizations.

Here, the obvious counterpoint is that Julia should outperform Python. However (I can’t speak for Knet), Flux is consistently slower than both PyTorch and TensorFlow on training and inference on common datasets (e.g. ImageNet).

Let me be clear that none of these points are an indictment of Julia the language, its potential or the community. The reality is that it took a lot of developer hours to get the Python DL frameworks to where they are now and the Julia ML/DL ecosystem hasn’t yet had nearly as much time poured into it. Will we get there? I hope so! Are we there now? Probably not.

[Next page](https://discourse.julialang.org/t/why-is-python-not-julia-still-used-for-most-state-of-the-art-ai-research/45896.md?page=2)
