# ArrayFire and Flux

**URL:** <https://discourse.julialang.org/t/arrayfire-and-flux/42907>\
**Category:** GPU\
**Tags:** question, arrayfire, flux\
**Created:** [July 11, 2020, 8:19pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907 "2020-07-11T20:19:01Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ian\_Slagle](https://avatars.discourse-cdn.com/v4/letter/i/b5a626/32.png) [@Ian\_Slagle](https://discourse.julialang.org/u/Ian_Slagle)\
**Post date:** [July 11, 2020, 8:19pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/1 "2020-07-11T20:19:01Z")

</div>

In part because of having an AMD GPU, the only package I can get to work with it is ArrayFire.jl.

```julia
ArrayFire v3.7.0 (OpenCL, 64-bit Linux, build a4485443)
[0] AMD: Ellesmere, 7999 MB
-1- INTEL: AMD Ryzen 7 2700X Eight-Core Processor , 32178 MB

```

Generally, the benchmarks for standard matrix operations look considerably better. The next step for me was to try to utilize my GPU with Flux.jl. I did this by a relatively naive approach, with my understanding of how CUDA interfaces with Flux. From my understanding, this meant converting the arrays of the network into ArrayFire arrays (which requires the arrays be converted to regular, untracked arrays I think), which is exactly what I did:

```julia
model = mapleaves(AFArray, mapleaves(Tracker.data, Chain(
  Dense(24, 24, σ),
  Dense(24, 24),
  softmax
)))

```

It’s a simple, small network that I’m just using for benchmarking, but I saw similar results across a number of different sizes and numbers of layers. In order for this to work, I had to extend the AFArray function for a few types, which I tried to do in as simple a way as possible.

```julia
ArrayFire.AFArray(func::typeof(σ)) = σ
ArrayFire.AFArray(func::typeof(identity)) = identity
ArrayFire.AFArray(func::typeof(softmax)) = softmax

```

And this worked, as in, it evaluated without error. Unfortunately, the GPU version performed much more slowly (0.048556 vs. 0.000040 seconds) for a single evaluation. Is there something that I’m doing wrong here or could be sped up?

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 11, 2020, 10:32pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/2 "2020-07-11T22:32:52Z")

</div>

sorry I don’t have the answer for you, I am a Linux and thus an AMD user as well, esp with their CPUs completely crushing Intel in the past few years I feel like they deserve more love in Julia community (for how terribly Nvidia treats FOSS ocmmunity).

Unfortunately devs don’t have/use AMD GPUs so the progress is slow, I honestly want to crowd source some money to buys some AMD GPUs for a few interested devs to make AMDGPU just work…

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [July 12, 2020, 12:24am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/3 "2020-07-12T00:24:38Z")

</div>

I believe that good progress is being made with AMD GPUs. cc @jpsamaroo

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [July 12, 2020, 12:27am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/4 "2020-07-12T00:27:24Z")

</div>

Yup, see [https://github.com/JuliaGPU/AMDGPU.jl/](https://github.com/JuliaGPU/AMDGPU.jl/). @jpsamaroo’s work there is funded IIRC, so it’s been progressing quite quickly!

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [July 12, 2020, 12:39am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/5 "2020-07-12T00:39:27Z")

</div>

Are you able to upgrade to a more recent version of Flux? `Tracker` and `mapleaves` have long since been deprecated in favour of Zygote and [fmap](https://github.com/JuliaGPU/AMDGPU.jl/).

c.f.

```julia
...
using AMDGPU

model = Chain(
  Dense(24, 24, σ),
  Dense(24, 24),
  softmax
)

m_gpu = fmap(HSAArray, m)
...

```

(not tested since I don’t have a compatible GPU, but given how close the CUDA and AMDGPU APIs are I suspect it’s not too far off)

> [@Ian\_Slagle](#):
>
> Unfortunately, the GPU version performed much more slowly (0.048556 vs. 0.000040 seconds) for a single evaluation. Is there something that I’m doing wrong here or could be sped up?

Unless you’re using a batch size in the high hundreds or thousands, the cost of multiplying 24xN matrices is going to be completely outweighed by communication and data transfer overhead with the GPU. These overheads should be amortized assuming your actual network is quite a bit larger.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 12, 2020, 12:56am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/6 "2020-07-12T00:56:43Z")

</div>

the problem at the moment is that I can’t even ~~install~~ build AMDGPU, since the dependence is even more non-trivial than CUDA

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [July 12, 2020, 1:09am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/7 "2020-07-12T01:09:38Z")

</div>

Are you able to run anything else that uses ROCm? If so, I would 100% open an issue. The maintainers are usually very responsive and equally patient 🙂

---

<div class="post-metadata">

**Author:** ![Ian\_Slagle](https://avatars.discourse-cdn.com/v4/letter/i/b5a626/32.png) [@Ian\_Slagle](https://discourse.julialang.org/u/Ian_Slagle)\
**Post date:** [July 12, 2020, 1:22am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/8 "2020-07-12T01:22:45Z")

</div>

> [@ToucheSir](#):
>
> Unless you’re using a batch size in the high hundreds or thousands, the cost of multiplying 24xN matrices is going to be completely outweighed by communication and data transfer overhead with the GPU. These overheads should be amortized assuming your actual network is quite a bit larger.

I was wondering about this. How would I ensure that the batching is working properly (I’ve tried them in a loop with similar performance unfortunately)? Is that through ArrayFire (or whatever GPU library)?

> [@ToucheSir](#):
>
> Are you able to upgrade to a more recent version of Flux? `Tracker` and `mapleaves` have long since been deprecated in favour of Zygote and [fmap](https://github.com/JuliaGPU/AMDGPU.jl/).
> 
> c.f.

I’ll give it a try thanks!

---

<div class="post-metadata">

**Author:** ![Ian\_Slagle](https://avatars.discourse-cdn.com/v4/letter/i/b5a626/32.png) [@Ian\_Slagle](https://discourse.julialang.org/u/Ian_Slagle)\
**Post date:** [July 12, 2020, 3:24am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/9 "2020-07-12T03:24:47Z")

</div>

> [@ToucheSir](#):
>
> Yup, see [https://github.com/JuliaGPU/AMDGPU.jl/](https://github.com/JuliaGPU/AMDGPU.jl/). @jpsamaroo’s work there is funded IIRC, so it’s been progressing quite quickly!

I definitely need to give this a try. It looks like the first release has happened since I last tried to get this package working on my computer.

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [July 12, 2020, 5:34am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/10 "2020-07-12T05:34:36Z")

</div>

Batching is done through Flux. I believe the last dimension is the batch dimension by convention.

---

<div class="post-metadata">

**Author:** ![Ian\_Slagle](https://avatars.discourse-cdn.com/v4/letter/i/b5a626/32.png) [@Ian\_Slagle](https://discourse.julialang.org/u/Ian_Slagle)\
**Post date:** [July 12, 2020, 7:03pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/11 "2020-07-12T19:03:01Z")

</div>

> [@ToucheSir](#):
>
> Are you able to upgrade to a more recent version of Flux? `Tracker` and `mapleaves` have long since been deprecated in favour of Zygote and [fmap](https://github.com/JuliaGPU/AMDGPU.jl/).
> 
> c.f.
> 
> ```julia
> ...
> using AMDGPU
> 
> model = Chain(
> Dense(24, 24, σ),
> Dense(24, 24),
> softmax
> )
> 
> m_gpu = fmap(HSAArray, m)
> ...
> 
> ```

Using the latest released (in Pkg) version of Flux like this yields this error:

```julia
UndefVarError: fmap not defined

```

Trying to install the master branch of Flux yields this strange error (considering AMDGPU is only on 0.1.0):

```julia
Unsatisfiable requirements detected for package GPUArrays [0c68f7d7]:
 GPUArrays [0c68f7d7] log:
 ├─possible versions are: [0.3.0-0.3.4, 0.4.0-0.4.2, 0.5.0, 0.6.0-0.6.1, 0.7.0-0.7.2, 1.0.0-1.0.4, 2.0.0-2.0.1, 3.0.0-3.0.1, 3.1.0, 3.2.0, 3.3.0, 3.4.0-3.4.1, 4.0.0] or uninstalled
 ├─restricted by compatibility requirements with AMDGPU [21141c5a] to versions: 2.0.0-2.0.1
 │ └─AMDGPU [21141c5a] log:
 │ ├─possible versions are: 0.1.0 or uninstalled
 │ └─restricted to versions * by an explicit requirement, leaving only versions 0.1.0
 └─restricted by compatibility requirements with CUDA [052768ef] to versions: 4.0.0 — no versions left
   └─CUDA [052768ef] log:
     ├─possible versions are: [0.1.0, 1.0.0-1.0.2, 1.1.0] or uninstalled
     └─restricted to versions 1 by Flux [587475ba], leaving only versions [1.0.0-1.0.2, 1.1.0]
       └─Flux [587475ba] log:
         ├─possible versions are: 0.11.0 or uninstalled
         └─Flux [587475ba] is fixed to version 0.11.0

```

---

<div class="post-metadata">

**Author:** ![Ian\_Slagle](https://avatars.discourse-cdn.com/v4/letter/i/b5a626/32.png) [@Ian\_Slagle](https://discourse.julialang.org/u/Ian_Slagle)\
**Post date:** [July 12, 2020, 11:10pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/12 "2020-07-12T23:10:04Z")

</div>

Update: I’ve installed AMDGPU.jl (and rocblas, rocrand, rocsparse, rocalution, rocfft, and MIOpen though not sure how necessary they are for this) per @ToucheSir and @dpsanders ’ suggestions . I then interfaced it with Flux in the exact same way as I did with ArrayFire:

```julia
AMDGPU.HSAArray(func::typeof(σ)) = σ
AMDGPU.HSAArray(func::typeof(identity)) = identity
AMDGPU.HSAArray(func::typeof(softmax)) = softmax
model2 = mapleaves(HSAArray, mapleaves(Tracker.data, Chain(
  Dense(24, 24, σ),
  Dense(24, 24),
  softmax
)))

```

For this small network, the GPU version did outperform the CPU (though only by a little: 1.498 μs vs. 1.788 μs) for evaluation of the neural network.

However, trying this with a larger network like this:

```julia
Chain(
  Dense(604, 400, σ),
  Dense(400, 300),
  Dense(300, 197),
  Dense(197, 197),
  Dense(197, 197),
  Dense(197, 110),
  Dense(110, 20),
  softmax
)

```

did not see significant performance improvments.

CPU:

```julia
BenchmarkTools.Trial: 
  memory estimate: 23.84 KiB
  allocs estimate: 117
  --------------
  minimum time: 53.960 μs (0.00% GC)
  median time: 62.500 μs (0.00% GC)
  mean time: 65.319 μs (2.06% GC)
  maximum time: 3.030 ms (96.96% GC)
  --------------
  samples: 10000
  evals/sample: 1

```

GPU:

```julia
BenchmarkTools.Trial: 
  memory estimate: 20.89 KiB
  allocs estimate: 22
  --------------
  minimum time: 201.469 μs (0.00% GC)
  median time: 206.269 μs (0.00% GC)
  mean time: 207.926 μs (0.38% GC)
  maximum time: 2.330 ms (85.28% GC)
  --------------
  samples: 10000
  evals/sample: 1

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [July 13, 2020, 3:11am UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/13 "2020-07-13T03:11:11Z")

</div>

> [@Ian\_Slagle](#):
>
> did not see significant performance improvments

instead, significant slow-doen

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [July 13, 2020, 9:19pm UTC](https://discourse.julialang.org/t/arrayfire-and-flux/42907/14 "2020-07-13T21:19:21Z")

</div>

It would be great if someone could test this with a CUDA GPU to compare. However, as stated, you may not get competitive performance over the CPU without using very expensive layers like convolutions.

Alternatively, it’s equally as likely that AMDGPU.jl is at fault for being slow. `HSAArray` isn’t really optimized for anything, and it’s slated to be merged with the `ROCArray`, which is supposed to be the more performant and featureful array type (we mostly just need the `HSAArray` for tests). Both array types currently do very bad things as well, such as falling back to running operations on the CPU silently (which CuArrays explicitly avoids, and we will too soon). Much of this will be fixed in the next few months, and I’ll have better news 🙂
