# GSoC 17' Proposal | Enabling Julia to target GPUs through Polly and imrpove code using run time information

**URL:** <https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872>\
**Category:** Community\
**Tags:** announcement\
**Created:** [March 25, 2017, 6:37am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872 "2017-03-25T06:37:09Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [March 25, 2017, 6:37am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/1 "2017-03-25T06:37:09Z")

</div>

Hello All,

I’m Sanjay Srivallabh, a final year undergraduate student of BITS-Pilani, Hyderabad Campus in India. I’m currently pursuing a semester long dissertation in Polyhedral Compilation under Dr. Ramakrishna at IIT-Hyderabad, India.

I’d like to take GSoC as learning opportunity and make a proposal to,

- Enable Julia to run on GPUs through Polly
- Better optimise code using run-time information
- Enable Polly to choose between sending code to the GPU or CPU for best results.

I’ll be sharing the link to my proposal in a while. 'Looking forward to your feedback and suggestions.

Thank You,  
Sanjay

---

<div class="post-metadata">

**Author:** ![MikeInnes](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikeinnes/32/3656_2.png) [@MikeInnes](https://discourse.julialang.org/u/MikeInnes)\
**Post date:** [March 27, 2017, 3:46pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/2 "2017-03-27T15:46:04Z")

</div>

Have you checked out CUDAnative.jl? Would this be an alternative to that?

---

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [March 28, 2017, 4:26am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/3 "2017-03-28T04:26:46Z")

</div>

Hello @MikeInnes,

CUDANative.jl requires you to understand CUDA syntax and the GPU execution model to write a function that runs on GPUs. Polly on the the other can turn any Julia code that conforms to certain specifications to NVPTX code. It’s an auto-parallelization pass in LLVM.

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [March 28, 2017, 8:32pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/4 "2017-03-28T20:32:26Z")

</div>

How will you handle the down and uploads to the GPU and automatically decide if latency + down/upload penalty are worth this optimization?  
Correct me if I’m wrong, but it seems like CUDAnative and the Polly approach are orthogonal. I’m guessing, that with Polly you only decide what code you _could_ execute on the GPU. But then you still need to take the LLVM code, compile it, link it and execute the GPU kernel, which is what CUDAnative can do!  
Or is this all in included in Polly already? That’d be pretty magical 🙂

---

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [March 28, 2017, 8:56pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/5 "2017-03-28T20:56:00Z")

</div>

> [@sdanisch](#):
>
> How will you handle the down and uploads to the GPU and automatically decide if latency + down/upload penalty are worth this optimization?

Polly currently uses a simple cost model that can decide this.

> [@sdanisch](#):
>
> Correct me if I’m wrong, but it seems like CUDAnative and the Polly approach are orthogonal.

Yes, they are orthogonal.

> [@sdanisch](#):
>
> Or is this all in included in Polly already?

Yes they are ! It already works with clang, have a look at [this](http://polly.llvm.org/documentation/gpgpucodegen.html)

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [March 28, 2017, 9:06pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/6 "2017-03-28T21:06:51Z")

</div>

I see. Not sure how this will apply to Julia. Also, one big problem isn’t solved yet: `Determine where to place the data.`.  
Probably an easier scope would be to use polly to optimize kernels which are written using for loops over CuArrays. So you could implement that as a pass inside CUDAnative.  
CC @maleadt, does this make sense?

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [March 29, 2017, 10:06am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/7 "2017-03-29T10:06:18Z")

</div>

The above GSOC project lists `Code Generation For Host` as only 50% done, and it also doesn’t seem to be mainlined? If that is the case, I would agree with @sdanisch that using the parts of Polly that already work (presumable recognizing parallelize loops, transforming them for optimal GPU execution, maybe some data placement, etc) either as a pass in CUDAnative, or as a new package building on top of CUDAdrv (or GPUArrays for a vendor-neutral alternative) might be a better choice.

Another reason it might be interesting to implement this as part of CUDAnative (or similar) is that you would obviously need to stick to a subset of Julia which Polly can analyze, and that can be executed on GPUs. CUDAnative is already doing exactly that, so it would be a shame if we’d introduce yet another new Julia’ → GPU compilation path.

Either way, looking forward to your proposal!

---

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [March 29, 2017, 5:17pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/8 "2017-03-29T17:17:19Z")

</div>

> [@maleadt](#):
>
> The above GSOC project lists Code Generation For Host as only 50% done, and it also doesn’t seem to be mainlined?

@Tobias_Grosser1 Could you please clarify what “Code Generation For Host as only 50% done” means ? Does it still reflect the current state of the project ?

@maleadt Polly generates NVPTX code and stores it as a string within LLVM-IR. It then inserts calls to a [runtime](https://github.com/llvm-mirror/polly/tree/master/tools/GPURuntime) to handle data transfers and launch the kernel. Also, GSoC project has been proposed to extend Polly to generate SPIR-V code for GPUs that don’t support NVPTX.

From what I understand from CUDANative’s repo, it lets you write functions meant just for the GPU in a syntax similar to that of CUDA kernels in Julia. Polly works at the higher level, turning general (and suitable) Julia code to NVPTX. So, I’m not sure if,

> [@sdanisch](#):
>
> Probably an easier scope would be to use polly to optimize kernels which are written using for loops over CuArrays.

is possible. @Tobias_Grosser1 Could you please share your thoughts on this ?

---

<div class="post-metadata">

**Author:** ![Tobias\_Grosser1](https://avatars.discourse-cdn.com/v4/letter/t/258eb7/32.png) [@Tobias\_Grosser1](https://discourse.julialang.org/u/Tobias_Grosser1)\
**Post date:** [March 29, 2017, 8:29pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/9 "2017-03-29T20:29:43Z")

</div>

@sanyam: The website at [Polly - GPGPU Code Generation](http://polly.llvm.org/documentation/gpgpucodegen.html) is completely outdated. It should be removed and replaced with actual documentation for Polly-ACC. If you are interested about performance numbers read (or at least skim) this paper: [http://grosser.es/bibliography/grosser2016pollyacc.html](http://grosser.es/bibliography/grosser2016pollyacc.html)

So yes, we can do fully automatic GPU code generation with Polly-ACC for two SPEC benchmarks and a variety of computational kernels. There are a lot more opportunities, so I believe this is indeed something pretty exciting for Julia.

Also, yes, this works like magic, “fully automatically without any user interaction”. Similar to real-world magic, it has constraints in what can be done. Still, I believe it would be a great idea to get this to Julia.

---

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [March 31, 2017, 12:59pm UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/10 "2017-03-31T12:59:34Z")

</div>

Hello All,

Here’s the [link to the draft of my GSoC proposal](https://docs.google.com/document/d/1od1FRptFvQpNhct8bQtb31oLy-UkK2oMwiUQ7RNeKWc/edit?usp=sharing). Please comment on it and let me know your suggestions.

Thank You,  
Sanjay

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [April 3, 2017, 6:24am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/11 "2017-04-03T06:24:34Z")

</div>

> [@Tobias\_Grosser1](#):
>
> So yes, we can do fully automatic GPU code generation with Polly-ACC for two SPEC benchmarks and a variety of computational kernels. There are a lot more opportunities, so I believe this is indeed something pretty exciting for Julia.
> 
> Also, yes, this works like magic, “fully automatically without any user interaction”. Similar to real-world magic, it has constraints in what can be done. Still, I believe it would be a great idea to get this to Julia.

How does it deal with rich CPU objects? I can image a C `float*` getting detected and offloaded, but would the same apply to eg. a `jl_array_t*` (containing another data pointer & metadata)? That’s where I figured some manual work at the frontier between host & `@polly` annotated code would be necessary.

@singam-sanjay: to elaborate on my comment on your proposal, I think it would be better to avoid adding Polly-specific flow to the main codegen, instead trying to use or improve existing mechanisms for outlining compiler functionality, and implementing your Polly-specific functionality as part of a package. This has many advantages (maintainability of both Julia’s codegen and your work, lower barrier to contributing, you get to work in Julia instead of C++, etc). Although that functionality (CodegenParams, CodegenHooks) has been developed for CUDAnative.jl, it is meant to be generic, and extending those mechanisms is not much work (adding more hooks or params, slightly restructuring codegen).

---

<div class="post-metadata">

**Author:** ![singam-sanjay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/singam-sanjay/32/1721_2.png) [@singam-sanjay](https://discourse.julialang.org/u/singam-sanjay)\
**Post date:** [April 3, 2017, 9:56am UTC](https://discourse.julialang.org/t/gsoc-17-proposal-enabling-julia-to-target-gpus-through-polly-and-imrpove-code-using-run-time-information/2872/12 "2017-04-03T09:56:57Z")

</div>

Thanks for the suggestions and info @maleadt !!
