# Does Flux.jl layers make use of tensor cores in Nvidia GPUs?

**URL:** <https://discourse.julialang.org/t/does-flux-jl-layers-make-use-of-tensor-cores-in-nvidia-gpus/103249>\
**Category:** GPU\
**Tags:** question, gpu, cuda, flux\
**Created:** [August 27, 2023, 7:23am UTC](https://discourse.julialang.org/t/does-flux-jl-layers-make-use-of-tensor-cores-in-nvidia-gpus/103249 "2023-08-27T07:23:10Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![lepton01](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lepton01/32/48824_2.png) [@lepton01](https://discourse.julialang.org/u/lepton01)\
**Post date:** [August 27, 2023, 7:23am UTC](https://discourse.julialang.org/t/does-flux-jl-layers-make-use-of-tensor-cores-in-nvidia-gpus/103249/1 "2023-08-27T07:23:10Z")

</div>

I was searching for an answer in the Flux.jl, CUDA.jl, cuDNN.jl, but only found [[https://juliagpu.org/2020-10-02-cuda\_2.0/#low--and-mixed-precision-operations](https://juliagpu.org/2020-10-02-cuda_2.0/#low--and-mixed-precision-operations)] which talks about independent GPU operations using CUDA.jl.  
I have not found particular information about Flux.jl exploiting this technology.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [August 28, 2023, 7:31am UTC](https://discourse.julialang.org/t/does-flux-jl-layers-make-use-of-tensor-cores-in-nvidia-gpus/103249/2 "2023-08-28T07:31:18Z")

</div>

I don’t think Flux uses mixed-precision, so probably no. It is possible to configure CUDA.jl to use tensor cores more eagerly, at the expense of some precision, by starting Julia with fast math enabled or by calling `CUDA.math_mode!(CUDA.FAST_MATH)`, which will e.g. use TF32 when doing an F32xF32 matmul. Further speed-ups are possible by setting CUDA.jl’s math precision to `:BFloat16` or even `:Float16`. Ideally though, I guess Flux.jl would have an interface to use mixed-precision arithmetic.
