# Fill up and fill down rows

**URL:** <https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196>\
**Category:** New to Julia\
**Tags:** dataframes\
**Created:** [April 28, 2022, 9:41am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196 "2022-04-28T09:41:13Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 28, 2022, 9:41am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/1 "2022-04-28T09:41:13Z")

</div>

Hi i have a dataframe looks like this

```julia
df1 = DataFrame()
df1.id = sort!(repeat(1:3,5))
df1.a = [1,missing,2,3,missing,missing,2,3,4,5, 1,2,3,missing,5]

```

i want to fill the missing values in column a with the previous value of same id

i want a dataframe like this

```julia
df2 = DataFrame()
df2.id = sort!(repeat(1:3,5))
df2.a = [1,1,2,3,3,missing,2,3,4,5, 1,2,3,3,5]

```

can somebody help me to do this

---

<div class="post-metadata">

**Author:** ![feanor12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/feanor12/32/8212_2.png) [@feanor12](https://discourse.julialang.org/u/feanor12)\
**Post date:** [April 28, 2022, 10:24am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/2 "2022-04-28T10:24:40Z")

</div>

Here is a quite verbose way of doing it. 😉

```julia
for gdf in groupby(df1,:id)
  for row_idx in 2:nrow(gdf)
    if ismissing(gdf.a[row_idx])
      gdf.a[row_idx] = gdf.a[row_idx-1]
    end
  end
end

```

---

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 28, 2022, 10:35am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/3 "2022-04-28T10:35:05Z")

</div>

what kind of midification should i do , to fill the value with next value of same id

---

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 28, 2022, 10:36am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/4 "2022-04-28T10:36:45Z")

</div>

is this fine

```julia
for gdf in groupby(df1,:id)
  for row_idx in 1:nrow(gdf)-1
    if ismissing(gdf.a[row_idx])
      gdf.a[row_idx] = gdf.a[row_idx + 1]
    end
  end
end

```

---

<div class="post-metadata">

**Author:** ![feanor12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/feanor12/32/8212_2.png) [@feanor12](https://discourse.julialang.org/u/feanor12)\
**Post date:** [April 28, 2022, 10:38am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/6 "2022-04-28T10:38:13Z")

</div>

It looks ok.

---

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 28, 2022, 10:38am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/7 "2022-04-28T10:38:53Z")

</div>

Thanks, is there any other way of doing it ?

---

<div class="post-metadata">

**Author:** ![feanor12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/feanor12/32/8212_2.png) [@feanor12](https://discourse.julialang.org/u/feanor12)\
**Post date:** [April 28, 2022, 10:46am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/8 "2022-04-28T10:46:55Z")

</div>

Here is a one-liner, but I find it hard to comprehend.

```julia
combine(groupby(df1,:id),:a=>(x->[x[1],coalesce.(x[2:end],x[1:end-1])...])=>:a)

```

---

<div class="post-metadata">

**Author:** ![sai\_matcha](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sai_matcha/32/35801_2.png) [@sai\_matcha](https://discourse.julialang.org/u/sai_matcha)\
**Post date:** [April 28, 2022, 10:47am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/9 "2022-04-28T10:47:27Z")

</div>

Thanks

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 28, 2022, 10:53am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/10 "2022-04-28T10:53:55Z")

</div>

If you want to update `df1` in-place do:

```julia
using Impute
for sdf in groupby(df1,:id)
    sdf.a .= Impute.locf(sdf.a)
end

```

or

```julia
transform!(groupby(df1, :id), :a => Impute.locf => :a)

```

if you want a new data frame:

```julia
transform(groupby(df1, :id), :a => Impute.locf => :a)

```

---

<div class="post-metadata">

**Author:** ![monopolynomial](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/monopolynomial/32/34782_2.png) [@monopolynomial](https://discourse.julialang.org/u/monopolynomial)\
**Post date:** [May 1, 2022, 1:06am UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/11 "2022-05-01T01:06:51Z")

</div>

`InMemoryDatasets` package has `ffill` and `bfill` similar to `pandas` functions.

```julia
using InMemoryDatasets
ds=Dataset(df1)
modify(IMD.groupby(ds,:id),:a=>ffill!)

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [February 28, 2023, 9:01pm UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/12 "2023-02-28T21:01:40Z")

</div>

an idea taken from an old post of mine

```julia
df = DataFrame(dt1=[missing, 0.2, missing, missing, 1, missing, 5, 6],
                      dt2=[9, 0.3, missing, missing, 3, missing, 5, 6])
filldown(v)=accumulate((x,y)->coalesce(y,x), v,init=v[1])

transform(df,[:dt1,:dt2].=>filldown,renamecols=false)

fillup(v)=reverse(filldown(reverse(v)))

transform(df,[:dt2,:dt1].=>[filldown,fillup],renamecols=false)

```

---

<div class="post-metadata">

**Author:** ![lrnv](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lrnv/32/19373_2.png) [@lrnv](https://discourse.julialang.org/u/lrnv)\
**Post date:** [February 28, 2023, 9:46pm UTC](https://discourse.julialang.org/t/fill-up-and-fill-down-rows/80196/13 "2023-02-28T21:46:27Z")

</div>

If I may profit from this discussion to ask: is there any performance reasons not to use the “verbose” loopy version ?

I know that loops are usually easier on the compiler, but is this reasoning still true for DataFrames ?
