# Matrix multipcation broadcasting?

**URL:** <https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483>\
**Category:** New to Julia\
**Created:** [June 17, 2023, 5:11am UTC](https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483 "2023-06-17T05:11:16Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![zxygentoo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zxygentoo/32/50824_2.png) [@zxygentoo](https://discourse.julialang.org/u/zxygentoo)\
**Post date:** [June 17, 2023, 5:11am UTC](https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483/1 "2023-06-17T05:11:16Z")

</div>

Hi everyone,

I’m trying to do the nanoGPT tutorial by Andrej Karpathy and stuck at a little piece of code something like this:

```python
import troch

B, T, C = 32, 8, 65

x = torch.randn(B, T, C)

wei = torch.tril(torch.ones(T, T))
wei = wei / wei.sum(1, keepdim=True)

xbow2 = wei @ x

```

Everything above the last line is easy:

```julia
using Flux

B, T, C = 32, 8, 65

x = randn(B, T, C)

wei = tril(ones(T, T))
wei = wei ./ sum(wei, dims=2)

```

For the last line, PyTorch seems batch wei (T, T) → (B, T, T) then do a batch operation, how to achieve the same effect in Julia (`repeat` and `broadcast` seems not working, or I’m doing it wrong, `batched_mul`?)? Or anywhere to look things up? (my linear algebra is really rusty at this point)

thx~

Alex

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [June 17, 2023, 5:47am UTC](https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483/2 "2023-06-17T05:47:34Z")

</div>

It would be good if you could show a minimal example of what the result in python is, but maybe you’re looking for `kronecker`?

---

<div class="post-metadata">

**Author:** ![bertschi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bertschi/32/33462_2.png) [@bertschi](https://discourse.julialang.org/u/bertschi)\
**Post date:** [June 17, 2023, 6:59am UTC](https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483/3 "2023-06-17T06:59:25Z")

</div>

For a direct translation you will need to handle the batching yourself, e.g.

```julia
xbow2 = Compat.stack(wei * x[b, :, :] for b ∈ axes(x, 1); dims = 1)

```

The real issue is that Julia arrays are column-major whereas Torch is row-major. Thus, Flux has the Tensor dimensions reversed as compared to Torch, i.e., the batch dimension last instead of first, and batched matrix multiplication correctly works when you also translate the tensors:

```julia
x = randn(C, T, B)

xbow2 = batched_mul(x, wei') # Note: Transpose wei as it was constructed according to Torch conventions

```

---

<div class="post-metadata">

**Author:** ![zxygentoo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/zxygentoo/32/50824_2.png) [@zxygentoo](https://discourse.julialang.org/u/zxygentoo)\
**Post date:** [June 17, 2023, 9:33am UTC](https://discourse.julialang.org/t/matrix-multipcation-broadcasting/100483/4 "2023-06-17T09:33:22Z")

</div>

`The real issue is that Julia arrays are column-major whereas Torch is row-major.`

Yes. exactly this. I’ve noticed the difference in previous code but not yet wrap my head around it (some onehot encoding stuff, I was able to line things up as default in Julia and avoiding reshaping to direclty match torch).

Thank for pointing it out! The goal here is not to mimic torch’s behavior but to work the tutuorial out and learn something about transformer/Julia.

Thank you again.
