# Initializing Flux weights the same as PyTorch?

**URL:** https://discourse.julialang.org/t/initializing-flux-weights-the-same-as-pytorch/54660
**Category:** Machine Learning
**Created:** [February 4, 2021, 11:09pm UTC](https://discourse.julialang.org/t/initializing-flux-weights-the-same-as-pytorch/54660 "2021-02-04T23:09:40Z")
**Posts on this page:** 1
**Showing post:** 4

<div class="post-metadata">

### Author: ![DevJac](https://avatars.discourse-cdn.com/v4/letter/d/50afbb/32.png) [@DevJac](https://discourse.julialang.org/u/DevJac)
#### Post date: [February 5, 2021, 5:05am UTC](https://discourse.julialang.org/t/initializing-flux-weights-the-same-as-pytorch/54660/4 "2021-02-05T05:05:55Z")

</div>

I came up with this function to initialize the weights the same way PyTorch does:

```julia
function Linear(in, out, activation)
    Dense(in, out, activation,
          initW=(_dims...) -> Float32.((rand(out, in).-0.5).*(2/sqrt(in))),
          initb=(_dims...) -> Float32.((rand(out).-0.5).*(2/sqrt(in))))
end

```

At least, for PyTorch’s `Linear` layers that’s how it works. You can easily verify this by creating a PyTorch `Linear` layer and looking at the minimum and maximum weight and bias values.

---

_[View the full topic](https://discourse.julialang.org/t/initializing-flux-weights-the-same-as-pytorch/54660)._
