# Julia: how to fill na in data frame like as data frame in python

**URL:** <https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317>\
**Category:** Community\
**Created:** [April 13, 2018, 6:51am UTC](https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317 "2018-04-13T06:51:43Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![bu\_ka](https://avatars.discourse-cdn.com/v4/letter/b/76d3ee/32.png) [@bu\_ka](https://discourse.julialang.org/u/bu_ka)\
**Post date:** [April 13, 2018, 6:51am UTC](https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317/1 "2018-04-13T06:51:43Z")

</div>

how to use the Julia to do the following python code in data frame Julia.  
(1) filling Na with most popular values  
(2) filling Na with mean

MSZoning NA in pred. filling with most popular values

```julia
features['MSZoning'] = features['MSZoning'].fillna(features['MSZoning'].mode()[0])

# LotFrontage NA in all. I suppose NA means 0
features['LotFrontage'] = features['LotFrontage'].fillna(features['LotFrontage'].mean())

```

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [April 13, 2018, 7:22am UTC](https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317/2 "2018-04-13T07:22:25Z")

</div>

You can try this (the example uses mean replacement):

```julia
julia> df = DataFrame(x = [1, 2, 9, missing])
4×1 DataFrames.DataFrame
│ Row │ x │
├─────┼─────────┤
│ 1 │ 1 │
│ 2 │ 2 │
│ 3 │ 9 │
│ 4 │ missing │

julia> recode!(df[:x], missing => mean(skipmissing(df[:x])));

julia> df
4×1 DataFrames.DataFrame
│ Row │ x │
├─────┼───┤
│ 1 │ 1 │
│ 2 │ 2 │
│ 3 │ 9 │
│ 4 │ 4 │

```

If you do not want to do it in-place use `recode` instead of `recode!`.

`mode` function is available in StatsBase.jl package.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [April 13, 2018, 12:41pm UTC](https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317/3 "2018-04-13T12:41:38Z")

</div>

Since you replace missing values, `df[:x] = recode(df[:x], missing => mean(skipmissing(df[:x])))` is probably better than `recode!` since it will create a column which does not allow for missing values, which makes further operations more efficient.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [April 13, 2018, 2:47pm UTC](https://discourse.julialang.org/t/julia-how-to-fill-na-in-data-frame-like-as-data-frame-in-python/10317/4 "2018-04-13T14:47:36Z")

</div>

In `DataFramesMeta` you can use an `ifelse` function.

```julia
df = DataFrame(rand(10, 10))
@transform(df, y = ifelse.(:x1 .> .5, :x1, 100))

```
