# Knet vs MXNet for programmer new to ML

**URL:** <https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073>\
**Category:** Machine Learning\
**Tags:** knet\
**Created:** [March 30, 2018, 12:22pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073 "2018-03-30T12:22:19Z")\
**Posts on this page:** 6\
**Page:** 2

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [April 2, 2018, 11:24pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/21 "2018-04-02T23:24:20Z")

</div>

That paper (and the rest of the information bottleneck work) has generated a fair amount of controversy (see [this ICLR paper and the open reviews](https://openreview.net/pdf?id=ry_WPG-A-)). For example, the “compression phase” where the mutual information decreases with more training doesn’t happen when you use ReLU (instead of tanh).

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [April 3, 2018, 12:10am UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/22 "2018-04-03T00:10:20Z")

</div>

Interesting. One comment I would have is that despite the extensive discussion about it in the original paper, I would not have expected the fitting and compression to have distinct and identifiable phases in the general case, but while that’s an interesting detail, it hardly seems like a crucial one. Second, I’d certainly expect that there would be cases where it’s _possible_ to get good representations of the target variable _without_ compression but (perhaps?) the interesting cases are those in which it’s too difficult to achieve this as a practical matter. That said, the examples where they show that they get over-fitting in spite of compression seem extremely worrying for the whole information bottleneck picture, so that’s definitely very interesting.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [October 6, 2018, 12:28pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/23 "2018-10-06T12:28:39Z")

</div>

Hello

What is faster Knet, Flux or MXNet?  
Which one is easier?  
Which one has more limitations?  
Which one can deal with larger datasets?

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [October 6, 2018, 1:22pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/24 "2018-10-06T13:22:25Z")

</div>

> [@Juan](#):
>
> What is faster Knet, Flux or MXNet?

KNet, at least for now.

> [@Juan](#):
>
> Which one is easier?

Flux

> [@Juan](#):
>
> Which one has more limitations?

KNet and MXNet

> [@Juan](#):
>
> Which one can deal with larger datasets?

All utilize standard Julia tools for the data handling, so tie unless this is a speed question.

Flux generates code from Julia functions, the others have hardcoded CUDA kernels in the package. Kernels in the package are easier to modify and get optimized for a single purpose, giving at least the current state of (1), but means you can’t just build any GPU kernel you want (3). Flux’s design is a major plus because it flows like a standard small Julia package whereas the other two flow like a standard ML framework which always felt to me “language in language” (2).

---

<div class="post-metadata">

**Author:** ![rickhg12hs](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rickhg12hs/32/5926_2.png) [@rickhg12hs](https://discourse.julialang.org/u/rickhg12hs)\
**Post date:** [October 6, 2018, 1:44pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/25 "2018-10-06T13:44:39Z")

</div>

For CPU only, how do they rank speedwise?

---

<div class="post-metadata">

**Author:** ![DoktorMike](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/doktormike/32/2736_2.png) [@DoktorMike](https://discourse.julialang.org/u/DoktorMike)\
**Post date:** [October 6, 2018, 1:48pm UTC](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073/26 "2018-10-06T13:48:11Z")

</div>

For CPU only and if you intend to use convolution then Knet currently is very slow. MXnet is quite ok and Flux I haven’t tried yet but I expect it to have MXnet speed.

[Previous page](https://discourse.julialang.org/t/knet-vs-mxnet-for-programmer-new-to-ml/10073.md?page=1)
