# Identical Random number generation in a DataFrame based on row categories

**URL:** <https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247>\
**Category:** New to Julia\
**Tags:** dataframes, random\
**Created:** [December 17, 2021, 8:16am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247 "2021-12-17T08:16:16Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![CompulsoryCoffee](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/compulsorycoffee/32/30434_2.png) [@CompulsoryCoffee](https://discourse.julialang.org/u/CompulsoryCoffee)\
**Post date:** [December 17, 2021, 8:16am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/1 "2021-12-17T08:16:16Z")

</div>

Hi,

I am trying to generate random numbers in a DataFrame that would be identical for the same categories.

for example:

```julia
│ Row │ id │ type │ mean │ std │
│ │ String │ String │ Float64 │ Float64 │
├─────┼────────┼────────┼─────────┼─────────┤
│ 1 │ A │ typeA │ 0.5 │ 0.2 │
│ 2 │ B │ typeA │ 0.5 │ 0.2 │
│ 3 │ C │ typeB │ 0.3 │ 0.1 │
│ 4 │ D │ typeB │ 0.3 │ 0.1 │

```

so I can generate a random number for each row like so:

```julia
d.rng1 = rand.(Normal.(d.mean, d.std))

│ Row │ id │ type │ mean │ std │ rng1 │
│ │ String │ String │ Float64 │ Float64 │ Float64 │
├─────┼────────┼────────┼─────────┼─────────┼──────────┤
│ 1 │ A │ typeA │ 0.5 │ 0.2 │ 0.265455 │
│ 2 │ B │ typeA │ 0.5 │ 0.2 │ 0.59307 │
│ 3 │ C │ typeB │ 0.3 │ 0.1 │ 0.310257 │
│ 4 │ D │ typeB │ 0.3 │ 0.1 │ 0.305229 │

```

but how to generate something that would look like column rng2 below based on the similarity in the column type? Is there a (fast) way to do it without creating a subset?

```julia
│ Row │ id │ type │ mean │ std │ rng1 │ rng2 │
│ │ String │ String │ Float64 │ Float64 │ Float64 │ Float64 │
├─────┼────────┼────────┼─────────┼─────────┼──────────┼──────────┤
│ 1 │ A │ typeA │ 0.5 │ 0.2 │ 0.265455 │ 0.265455 │
│ 2 │ B │ typeA │ 0.5 │ 0.2 │ 0.59307 │ 0.265455 │
│ 3 │ C │ typeB │ 0.3 │ 0.1 │ 0.310257 │ 0.310257 │
│ 4 │ D │ typeB │ 0.3 │ 0.1 │ 0.305229 │ 0.310257 │

```

Thanks

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 17, 2021, 8:41am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/3 "2021-12-17T08:41:28Z")

</div>

Seed the rng using the hash of the common key

reproducable

```julia
julia> transform!(d, :type => ByRow(t->rand(MersenneTwister(hash(t)))) => :rng)
50×2 DataFrame
 Row │ type rng
     │ String Float64
─────┼───────────────────
   1 │ type0 0.878381
   2 │ type5 0.92784
   3 │ type6 0.625461
   4 │ type7 0.141122
   5 │ type8 0.847776
   6 │ type8 0.847776
   7 │ type8 0.847776
   8 │ type4 0.841166
   9 │ type4 0.841166
  10 │ type2 0.905406

```

random

```julia
julia> salt=round(Int, 10000rand()); transform!(d, :type => ByRow(t->rand(MersenneTwister(hash(t)+salt))) => :rng)
50×2 DataFrame
 Row │ type rng
     │ String Float64
─────┼───────────────────
   1 │ type0 0.105401
   2 │ type5 0.165727
   3 │ type6 0.261818
   4 │ type7 0.0470375
   5 │ type8 0.56809
   6 │ type8 0.56809
   7 │ type8 0.56809
   8 │ type4 0.399268
   9 │ type4 0.399268
  10 │ type2 0.533638

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [December 17, 2021, 8:59am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/4 "2021-12-17T08:59:09Z")

</div>

If I understand you correctly you just want to draw one random number per group (rather than per row) and have that show up in all rows of that group? If so this should do it:

```julia
julia> transform!(groupby(df, :type), :id => (x -> rand()) => :rng)
4×3 DataFrame
 Row │ id type rng      
     │ String Int64 Float64  
─────┼─────────────────────────
   1 │ A 1 0.678716
   2 │ B 1 0.678716
   3 │ C 2 0.71421
   4 │ D 2 0.71421

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 17, 2021, 10:51am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/5 "2021-12-17T10:51:04Z")

</div>

> [@nilshg](#):
>
> `transform!(groupby(df, :type), :id => (x -> rand()) => :rng)`

Yes, or:

```julia
transform!(groupby(df, :type), [] => rand => :rng)

```

which is a bit shorter to type.

---

<div class="post-metadata">

**Author:** ![CompulsoryCoffee](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/compulsorycoffee/32/30434_2.png) [@CompulsoryCoffee](https://discourse.julialang.org/u/CompulsoryCoffee)\
**Post date:** [December 17, 2021, 11:21am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/6 "2021-12-17T11:21:04Z")

</div>

And if I wanted to use the mean and std columns as

```julia
Normal.(d.mean, d.std)

```

How would you write down

```julia
transform!(groupby(df, :type), [] => rand => :rng)

```

?

Thanks a lot for the help

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [December 17, 2021, 11:29am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/7 "2021-12-17T11:29:39Z")

</div>

Oh wow, there’s always new things in the minilanguage to discover! I had tried

```julia
:id => rand => :rng

```

initially but that of course won’t work because it essentially calls `rand(x, length(x)` on each subgroup-vector `x`(so essentially samples a random `id` withing the group.

Is the use of `[]` as column selector documented somewhere?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [December 17, 2021, 11:33am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/8 "2021-12-17T11:33:39Z")

</div>

You can do

```julia
transform!(groupby(df, :type), [:mean, :std] => ((mean, std) -> rand(Normal(first(mean), first(std)))) => :rng)

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 17, 2021, 11:56am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/9 "2021-12-17T11:56:00Z")

</div>

> [@nilshg](#):
>
> Is the use of `[]` as column selector documented somewhere?

It is just a vector selector, just like e.g. `[:a, :b]` or any other column selector. Just that it is am empty vector.

If it feels unintuitive to you use:

```julia
transform!(groupby(df, :type), Cols() => rand => :rng)

```

and now I hope it is clear that `Cols()` selects no columns (I use `[]` as it is shorter to type 😄).

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [December 17, 2021, 11:57am UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/10 "2021-12-17T11:57:16Z")

</div>

Makes sense - I guess I never came across

```julia
julia> select(df, [])
0×0 DataFrame

```

(as I suppose it is indeed a rarely used selector…)

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [December 17, 2021, 12:17pm UTC](https://discourse.julialang.org/t/identical-random-number-generation-in-a-dataframe-based-on-row-categories/73247/11 "2021-12-17T12:17:03Z")

</div>

or just:

```julia
julia> df[:, []]
0×0 DataFrame

```

Note that the same works with `AbstractArray`s:

```julia
julia> x = [1, 2, 3]
3-element Vector{Int64}:
 1
 2
 3

julia> x[[]]
Int64[]

julia> x = [1 2; 3 4]
2×2 Matrix{Int64}:
 1 2
 3 4

julia> x[:, []]
2×0 Matrix{Int64}

```

so, as usual, you get what Julia Base ships.
