# \[ANN\] EvoTrees.jl: experimental GPU support for gradient boosting trees

**URL:** <https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726>\
**Category:** Package Announcements\
**Tags:** machine-learning\
**Created:** [August 11, 2020, 4:51am UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726 "2020-08-11T04:51:19Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [August 11, 2020, 4:51am UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726/1 "2020-08-11T04:51:19Z")

</div>

Starting with v0.5.0, it’s now possible to build GBT models on GPU, thanks to the underlying CUDA.jl package providing nimble tools to handle kernels. Speedups are modest compared to those observed with XGBoost’s gpu\_hist approach as there likely are remaining optimizations along the CPU-GPU traffic or more efficient kernels.

A model can be trained on gpu using:  
`fit_evotree_gpu(params1, X, Y)` instead of the usual `fit_evotree(params1, X, Y)`.  
For predictions: `predict_gpu(model, X_train)`

> **[GitHub - Evovest/EvoTrees.jl: Boosted trees in Julia](https://github.com/Evovest/EvoTrees.jl)**
>
> Boosted trees in Julia. Contribute to Evovest/EvoTrees.jl development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 11, 2020, 7:40am UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726/2 "2020-08-11T07:40:06Z")

</div>

do I have to convert X to CuArrays’ first?

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [August 11, 2020, 5:43pm UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726/3 "2020-08-11T17:43:35Z")

</div>

No, the input should still be regular Matrix. GPU operations are performed for the gradient update and histograms accumulation all happening in the background.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 12, 2020, 3:14am UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726/4 "2020-08-12T03:14:30Z")

</div>

Would it better to run GPU only if user has converted to CuArrays.jl then u can benchmark just the computation and not the data transfer?

---

<div class="post-metadata">

**Author:** ![jeremiedb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeremiedb/32/29150_2.png) [@jeremiedb](https://discourse.julialang.org/u/jeremiedb)\
**Post date:** [August 20, 2020, 10:14pm UTC](https://discourse.julialang.org/t/ann-evotrees-jl-experimental-gpu-support-for-gradient-boosting-trees/44726/5 "2020-08-20T22:14:40Z")

</div>

I think it would effectively makes sense to run straight on GPU if the input is a CuArray. However, as of now, as there’s not such equivalent to Rapids/CuDf in Julia as far as I know, I wouldn’t necessarily expect input data to be naturally brought in as a CuArray. Also, current conversation to CuArray within the algorithm routine is aligned with the xgboost/lightgbm approach where input is also converted to GPU within the training routine so the benchmark between EvoTrees and others is I think fair with the current approach. Anyhow, for what I’ve observed, this step doesn’t add much to the global training time.  
I think main bottleneck in current implementation is the back and forth between CPU and GPU as the histogram is built on GPU but the best split scan is done on CPU.
