# Fast CNN inference

**URL:** <https://discourse.julialang.org/t/fast-cnn-inference/51761>\
**Category:** GPU\
**Created:** [December 13, 2020, 1:31pm UTC](https://discourse.julialang.org/t/fast-cnn-inference/51761 "2020-12-13T13:31:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fabrice\_Rosay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabrice_rosay/32/15689_2.png) [@Fabrice\_Rosay](https://discourse.julialang.org/u/Fabrice_Rosay)\
**Post date:** [December 13, 2020, 1:31pm UTC](https://discourse.julialang.org/t/fast-cnn-inference/51761/1 "2020-12-13T13:31:12Z")

</div>

Hi,  
I’m currently using Knet and Flux for doing alphazero like calculations. As the botlleneck is the many inferences needed i was wondering if it would be possible to wrap TensorRT like [https://github.com/zerollzeng/tiny-tensorrt](https://github.com/zerollzeng/tiny-tensorrt) (the C++ part) and how hard it would be. If not is it plausible that specific CUDA kernels for the inference could bring some acceleration, perhaps using Float16 ?  
Thanks.  
To give an idea 90% of the time is spent by this function, in which 90% of the time is spent on π,v=m(KnetArray(batch)) where m is a residual network. The typical size input is 8x8x8x1600 (reversi, 200 games in parallel, batch size for parallel MCTS 8) and inference time is around 150 ms for 96 layers 10 block using GTX 1070:(

```julia
function (m::resnetwork)(x::Vector{GameEnv},squashing=1f0)
    l=size(x)[1]
    batch=zeros(Float32,(sizeInput...,l))

    @threads for k in 1:l
         @views decoder(x[k],batch[:,:,:,k])
    end
   π,v=m(KnetArray(batch))
   π=softmax(squashing .*π)
   Array(π),Array(v)
end

```

---

<div class="post-metadata">

**Author:** ![ToucheSir](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/touchesir/32/14411_2.png) [@ToucheSir](https://discourse.julialang.org/u/ToucheSir)\
**Post date:** [December 13, 2020, 7:32pm UTC](https://discourse.julialang.org/t/fast-cnn-inference/51761/2 "2020-12-13T19:32:21Z")

</div>

Have you tried just keeping everything on CPU? That input size isn’t terribly massive, so I wonder if the back-and-forth transfer is worth the latency.

---

<div class="post-metadata">

**Author:** ![Fabrice\_Rosay](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fabrice_rosay/32/15689_2.png) [@Fabrice\_Rosay](https://discourse.julialang.org/u/Fabrice_Rosay)\
**Post date:** [December 14, 2020, 7:50am UTC](https://discourse.julialang.org/t/fast-cnn-inference/51761/3 "2020-12-14T07:50:34Z")

</div>

I tried it is at least a hundred times slower. I don’t know of any reasonable implementation that does not use gpu, that is if you try anything bigger than tic tac toe.

---

<div class="post-metadata">

**Author:** ![maleadt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/maleadt/32/10097_2.png) [@maleadt](https://discourse.julialang.org/u/maleadt)\
**Post date:** [December 14, 2020, 8:01am UTC](https://discourse.julialang.org/t/fast-cnn-inference/51761/4 "2020-12-14T08:01:22Z")

</div>

Wrapping C++ is tricky, and TensorRT doesn’t seem to have an C API either, so that won’t be easy. Maybe you can PyCall the TensorRT Python bindings.
