# \`CUDA.quicksort\` is not faster than CPU radixsort

**URL:** <https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036>\
**Category:** GPU\
**Tags:** performance\
**Created:** [July 4, 2021, 12:36pm UTC](https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036 "2021-07-04T12:36:32Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [July 4, 2021, 12:36pm UTC](https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036/1 "2021-07-04T12:36:32Z")

</div>

Here’s my testing code

```julia-auto
using SortingAlgorithms

a32 = rand(Float32, 10_000_000)

@benchmark sort($a32, alg=RadixSort)

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/7/6/763498351386b8f97ff7e537451fec8ff14fd88e.png)

```julia-auto

using CUDA

cua32 = cu(a32)

@benchmark CUDA.@sync sort($cua32)

```

 ![image](https://global.discourse-cdn.com/julialang/original/3X/1/9/19680062929dd5595dd68e719401cbf44158f1d1.png)

So the on my GPU which is RTX2080. CUDA sorting is not faster than radixsort. YMMV, obviously.

## 100million elements compared

So increasing the size of 100 million elements, and as you

 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/8/588c578c9997b60bc3e959c1d95c4c217c1c24da.png)

 ![image](https://global.discourse-cdn.com/julialang/original/3X/6/9/693f6d253ea8c05eaf912c3e93227d05ffef727d.png)

## GPU radix sort needed?

The issue might be that the sort as implemented in CUDA is quicksort whereas a radix sort implementation on GPU would be blazing fast. I was trying something along these lines but I think I got stuck somewhere with the older CUDA.

## How to load Float64 onto GPU?

```julia-auto
a64 = rand(10_000_000)

cua64 = cu(a64)

```

We can see that `cua6` is `CuArray{Float32}`. Anyway, GPUs like 32bit floats and have much better performance, so I can see why `cu` does the implicit conversion. But it would be nice to have an option to copy the `Float64` direct and `cua64 = Float64.(cu(a64)` has obvious issues.

Anyhow, I tested it,

 ![image](https://global.discourse-cdn.com/julialang/original/3X/1/7/17c6cfb456e5cf0d66caa0aa553ab21afd0c2073.png)

So I think it’s fair to say GPU sort as of now is no faster than he fastest available CPU sort. Also CPU sort tends to have access to more RAM vs GPU (I think), so CPU can handle more dataset anyway.

```julia-auto
sort(a64, alg=RadixSort) == collect(sort(cua64))

```

---

<div class="post-metadata">

**Author:** ![findmyway](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/findmyway/32/4946_2.png) [@findmyway](https://discourse.julialang.org/u/findmyway)\
**Post date:** [July 4, 2021, 3:18pm UTC](https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036/2 "2021-07-04T15:18:10Z")

</div>

> [@xiaodai](#):
>
> But it would be nice to have an option to copy the `Float64` direct and `cua64 = Float64.(cu(a64)` has obvious issues.

Try `CuArray{Float64}(a64)` instead.

---

<div class="post-metadata">

**Author:** ![eliassno](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eliassno/32/18917_2.png) [@eliassno](https://discourse.julialang.org/u/eliassno)\
**Post date:** [July 4, 2021, 3:28pm UTC](https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036/3 "2021-07-04T15:28:48Z")

</div>

I think `CuArray(a64)` is sufficient.

---

<div class="post-metadata">

**Author:** ![calebwin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/calebwin/32/14432_2.png) [@calebwin](https://discourse.julialang.org/u/calebwin)\
**Post date:** [July 4, 2021, 7:53pm UTC](https://discourse.julialang.org/t/cuda-quicksort-is-not-faster-than-cpu-radixsort/64036/4 "2021-07-04T19:53:08Z")

</div>

It’s still a useful operation though. If you are doing computation on the GPU that actually _is_ faster than CPU, then it’s ideal to keep your data on the GPU and continue operating on it where it is instead of moving it to the CPU and doing the sort and then moving it back.

So if your workload is just a single sort - sure stick with CPU for now. But if you’re doing more than just sorting, then it can be quite useful to be able to sort your data without expensive transfers between CPU and GPU.
