# Tensor regression models in Julia

**URL:** <https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601>\
**Category:** Machine Learning\
**Created:** [June 12, 2018, 5:55am UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601 "2018-06-12T05:55:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![outlace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/outlace/32/2216_2.png) [@outlace](https://discourse.julialang.org/u/outlace)\
**Post date:** [June 12, 2018, 5:55am UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/1 "2018-06-12T05:55:55Z")

</div>

I’m trying to create a tensor regression model in Julia. I first tried in PyTorch but it was becoming a hassle so thought I’d give Julia a try, and its looking like Julia can’t do it at all (at least in any efficient way).

A tensor regression is just a series of tensor contractions (“generalized matrix multiplications”), where the tensors (multi-dimensional arrays) act as trainable parameters. I want to train these parameter tensors using gradient descent as I would with a normal neural network (which is a series of matrix-matrix multiplications with interspersed non-linear functions). The problem is TensorOperations.jl, the major library that supports tensor contractions, does not work with either of Julia’s main ML libraries, Knet.jl and Flux.jl. Both of the latter libraries secretly convert your Array types into special trackable types so they can keep track of gradients, but then TensorOperations.jl can’t handle those types.

I tried Einsum.jl and it technically works with Flux.jl, but only with outrageously bad performance. It took literally 2 minutes to do a very small tensor contraction with Flux tracked arrays (takes 1.3 seconds with normal arrays).

I’m picking up Julia again after trying it back at v0.3 so I may just be missing something here. Any help appreciated.

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [June 12, 2018, 9:19am UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/2 "2018-06-12T09:19:16Z")

</div>

How complicated are the contractions you need? If you can re-write them with `reshape` and `permutedims` then Flux (and probably other options) should work well. Here’s A\_{i,j,k} B\_k = C\_{ij}:

```julia
using Flux
A = param(rand(2,2,7))
B = param(rand(7))
reshape(A, 4,7) * B |> C->reshape(C, 2,2)

```

---

<div class="post-metadata">

**Author:** ![improbable22](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/improbable22/32/5464_2.png) [@improbable22](https://discourse.julialang.org/u/improbable22)\
**Post date:** [June 12, 2018, 9:50am UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/3 "2018-06-12T09:50:29Z")

</div>

After reading the [readme](https://github.com/Jutho/TensorOperations.jl#building-blocks): TensorOperations seems to reduce everything to three functions `add!`, `trace!`, `contract!`.

It seems like it would not be very hard to write fallback versions of these functions in terms of reshape etc. as above. Or in fact to just work out a gradient for each of these, and then provide this to `Flux.back`.

---

<div class="post-metadata">

**Author:** ![outlace](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/outlace/32/2216_2.png) [@outlace](https://discourse.julialang.org/u/outlace)\
**Post date:** [June 12, 2018, 5:33pm UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/4 "2018-06-12T17:33:59Z")

</div>

I’m trying to do this contraction:

C\_{s,k} = X\_{s,h}A\_{h,i,j}B\_{i,j,k}

where X is some input tensor of data and the tensors A and B are the trainable parameter tensors. In particular I’m trying to do a tensor regression with MNIST data, so X is [batch size, 784] and the result is [batch size, 10] for each of digit classes.

I’ll try what you suggested.

---

<div class="post-metadata">

**Author:** ![jessebett](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jessebett/32/2006_2.png) [@jessebett](https://discourse.julialang.org/u/jessebett)\
**Post date:** [June 12, 2018, 6:15pm UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/5 "2018-06-12T18:15:17Z")

</div>

I agree with @improbable22 that these gradients should definitely be added as a method to Flux.back.

---

<div class="post-metadata">

**Author:** ![LaurentPlagne](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/laurentplagne/32/10103_2.png) [@LaurentPlagne](https://discourse.julialang.org/u/LaurentPlagne)\
**Post date:** [June 12, 2018, 6:21pm UTC](https://discourse.julialang.org/t/tensor-regression-models-in-julia/11601/6 "2018-06-12T18:21:27Z")

</div>

Slightly off topic but it reminds me that I wonder if the work related in this [article](https://arxiv.org/abs/1802.04730) could be efficiently developed in Julia…
