# How to permute the rows of a DataFrame in-place efficiently?

**URL:** <https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825>\
**Category:** Data\
**Tags:** performance, dataframes\
**Created:** [February 5, 2018, 2:16am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825 "2018-02-05T02:16:17Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [February 5, 2018, 2:16am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/1 "2018-02-05T02:16:17Z")

</div>

I have a vector of row numbers and I want to use it to permute a DataFrame’s rows. Here is an MVE

```julia
using StatsBase
df = DataFrame(a = rand(1_000_000))
r=sample(1:size(df,1), size(df,1), replace=false)
@time df = df[r,:]

```

I think the above creates a DataFrame and then assigns it to `df`. Is there a way to re-assign the rows in place so minimal extra memory is allocated?

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [February 5, 2018, 4:37am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/2 "2018-02-05T04:37:24Z")

</div>

What exactly are you trying to accomplish? Usually for group operations a `groupby` à la split and apply works pretty good and handles memory quite well.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [February 5, 2018, 5:36am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/3 "2018-02-05T05:36:55Z")

</div>

as i have shown. i want to randomize all the rows as efficiently as I can

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [February 5, 2018, 5:48am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/4 "2018-02-05T05:48:16Z")

</div>

```julia
using StatsBase, DataFrames
df = DataFrame(a = rand(1_000_000))
@time df = df[sample(1:size(df,1), size(df,1), replace=false),:];
@time df = df[shuffle(1:size(df, 1)),:];

```

`shuffle` is about a three times more efficient and about 2/3 less memory intensive (also drops `StatsBase` dependency)

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [February 5, 2018, 5:57am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/5 "2018-02-05T05:57:49Z")

</div>

> [@Nosferican](#):
>
> df = df[shuffle(1:size(df, 1)),:]

is this line the best possible? Does it not create a new DataFrame?

anyway i got an idea now

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 5, 2018, 9:47am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/6 "2018-02-05T09:47:44Z")

</div>

You can simply iterate over columns and call `permute!` on them.

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [February 5, 2018, 7:31pm UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/7 "2018-02-05T19:31:51Z")

</div>

```julia
using DataFrames
srand(0)
df = DataFrame(a = rand(Int(1e6)));
function permute_df_bycol!(df::AbstractDataFrame)
    p = shuffle(1:size(df, 1))
    for (name, col) ∈ eachcol(df)
        permute!(col, p)
    end
end
function permute_df!(df::AbstractDataFrame)
    df[:,:] = df[shuffle(1:size(df, 1)),:]
    return
end
@time permute_df_bycol!(df) # 0.064590 seconds (10 allocations: 15.259 MiB)
@time permute_df!(df) # 0.021940 seconds (36 allocations: 15.261 MiB)

```

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [February 5, 2018, 7:35pm UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/8 "2018-02-05T19:35:54Z")

</div>

that was my idea too. But it’s slower? But there is potential to apply to all columns using threads.

Maybe it’s faster for multiple columns

---

<div class="post-metadata">

**Author:** ![Nosferican](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nosferican/32/9275_2.png) [@Nosferican](https://discourse.julialang.org/u/Nosferican)\
**Post date:** [February 6, 2018, 11:45am UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/9 "2018-02-06T11:45:27Z")

</div>

You could try it, but the more columns you have the more work it would have to do. Rearranging every row for the whole dataframe should be faster as it only has to apply the permutation once and It writes in-place. Column wise might be a last resort if memory is not enough.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [February 6, 2018, 12:23pm UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/10 "2018-02-06T12:23:05Z")

</div>

> [@Nosferican](#):
>
> Rearranging every row for the whole dataframe should be faster as it only has to apply the permutation once and It writes in-place. Column wise might be a last resort if memory is not enough.

I don’t think “apply the permutation only once” is correct here. Under the hood, indexing a data frame implies indexing repeatedly each column.

I suspect the `permute!`-based sorting of `DataFrame` is slower just because permuting a vector is slower than indexing it (i.e. allocating a new copy). The overhead related to `DataFrame` should be negligible here.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [February 6, 2018, 1:13pm UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/11 "2018-02-06T13:13:28Z")

</div>

Why not use a `view`?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [February 20, 2018, 11:29pm UTC](https://discourse.julialang.org/t/how-to-permute-the-rows-of-a-dataframe-in-place-efficiently/8825/12 "2018-02-20T23:29:27Z")

</div>

Say, the computation I want to perform with the permuted dataframe would be faster if the all the columns are permuted as well. This is for “cache-efficiency”, as the next part of the program requires me to go through the vector several times in order.
