# Dummy Encoding(One hot encoding) from PooledDataArray

**URL:** <https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167>\
**Category:** General Usage\
**Tags:** question\
**Created:** [June 9, 2017, 2:12am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167 "2017-06-09T02:12:21Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 2:12am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/1 "2017-06-09T02:12:21Z")

</div>

I would like to know how to convert Pooled Data array into 0/1 columns similar to Sklearn OneHotEncoder in python. Following is the data frame which i am working with in which Country and Purchased are pooled data array.

```julia
10×4 DataFrames.DataFrame
│ Row │ Country │ Age │ Salary │ Purchased │
├─────┼───────────┼─────┼────────┼───────────┤
│ 1 │ "France" │ 44 │ 72000 │ "No" │
│ 2 │ "Spain" │ 27 │ 48000 │ "Yes" │
│ 3 │ "Germany" │ 30 │ 54000 │ "No" │
│ 4 │ "Spain" │ 38 │ 61000 │ "No" │
│ 5 │ "Germany" │ 40 │ NA │ "Yes" │
│ 6 │ "France" │ 35 │ 58000 │ "Yes" │
│ 7 │ "Spain" │ NA │ 52000 │ "No" │
│ 8 │ "France" │ 48 │ 79000 │ "Yes" │
│ 9 │ "Germany" │ 50 │ 83000 │ "No" │
│ 10 │ "France" │ 37 │ 67000 │ "Yes" │

```

`> pool!(data1csv, [:Country, :Purchased])`

Kindly let me know how to go about converting pooled data array into dummy encoded columns

Thank You

---

<div class="post-metadata">

**Author:** ![jkbest2](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jkbest2/32/7350_2.png) [@jkbest2](https://discourse.julialang.org/u/jkbest2)\
**Post date:** [June 9, 2017, 4:34am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/2 "2017-06-09T04:34:08Z")

</div>

I’m not familiar with `OneHotEncoder`, but a `ModelMatrix` from the `DataFrames` package is probably what you’re looking for. There is some documentation here:

[https://juliastats.github.io/DataFrames.jl/stable/man/formulas/](https://juliastats.github.io/DataFrames.jl/stable/man/formulas/)

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 5:23am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/3 "2017-06-09T05:23:50Z")

</div>

Thank you for the response. I did go through the documentation and was able to create it with dummy encoding as below

> `>mm = ModelMatrix(ModelFrame(@formula(Purchased ~ Age + Salary + Country), data1csv, contrasts = Dict(:Purchased => DummyCoding(), :Country => DummyCoding())))`

 ![](https://global.discourse-cdn.com/julialang/original/3X/7/c/7c8ef37bc508115446067157a856bca0a6389d4a.png)

> A ModelFrame object is just a simple wrapper around a DataFrame. For modeling purposes, one generally wants to construct a ModelMatrix, which constructs a Matrix{Float64} that can be used directly to fit a statistical model:

Document does say that it can be used to fit statistical model. But i am not sure how to even access the modelmatrix so that i can apply normalization function to it. Tried to search the web for information but unable to find the further document related to this? Please point me in the right direction if you are aware of it?

Thank You.

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 6:45am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/4 "2017-06-09T06:45:51Z")

</div>

I was able to figure out how to get the matrix from model matrix. I just had to use **m** property from the object to get the matrix.  
`> mm.m`

```julia
10×5 Array{Float64,2}:
 1.0 44.0 72000.0 0.0 0.0
 1.0 27.0 48000.0 0.0 1.0
 1.0 30.0 54000.0 1.0 0.0
 1.0 38.0 61000.0 0.0 1.0
 1.0 40.0 63777.0 1.0 0.0
 1.0 35.0 58000.0 0.0 0.0
 1.0 38.0 52000.0 0.0 1.0
 1.0 48.0 79000.0 0.0 0.0
 1.0 50.0 83000.0 1.0 0.0
 1.0 37.0 67000.0 0.0 0.0

```

After i apply normalization function to Age and Salary. How do i get back the original fame. As **Country** Feature is Dummy Encoded. How to associate a observation to respective country(France, Germany, Spain).  
Is there a way to do it?

---

<div class="post-metadata">

**Author:** ![mwsohn](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mwsohn/32/2696_2.png) [@mwsohn](https://discourse.julialang.org/u/mwsohn)\
**Post date:** [June 9, 2017, 7:27am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/5 "2017-06-09T07:27:38Z")

</div>

The simplest way may be to write a function to `unpool` a PooledDataArray. The following `unpool` function returns a DataArray, which you can populate to your original dataframe: df[:unpooled] = unpool(df,:pooled).

```julia
function unpool(df::DataFrame,varname::Symbol)
    if isa(df[varname],PooledDataArray) == false
        error(varname," is not a PooledDataArray")
    end

    da = DataArray(eltype(df[varname].pool),size(df,1))
    pool = df[varname].pool
    refs = df[varname].refs
    for i = 1:size(df,1)
        da[i] = refs[i] == 0 ? NA : pool[refs[i]]
    end
    return da
end

```

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [June 9, 2017, 9:47am UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/6 "2017-06-09T09:47:26Z")

</div>

It doesn’t seem you should used model matrices, as IIUC you want a data frame result rather than a matrix. A loop should be enough:

```julia
for c in unique(df[:Country])
    df[Symbol(c)] = df[:Country] .== c
end

```

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 3:12pm UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/7 "2017-06-09T15:12:33Z")

</div>

@mwsohn thank you. will try it out

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 3:19pm UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/8 "2017-06-09T15:19:56Z")

</div>

@nalimilan thank you for the response. You solution works perfectly. But if i need to convert true to 1 false to 0. Is there a efficient way to go about it as opposed to creating a new Float64/Int64 Column and deleting the Bool Column.  
I tried the following. As the new are columns Bool, I am unable to assign True to 1 and False to 0.

```julia
> for c in unique(df[:Country])
     df[Symbol(c)] = df[:Country] .== c
   
     for i in 1:size(df[Symbol(c)], 1)
       if df[i, Symbol(c)]
         df[i, Symbol(c)] = 1.0
       else
         df[i, Symbol(c)] = 0.0
       end
     end
end

```

If there is any alternative kindly suggest me.

Thank You.

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 4:30pm UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/9 "2017-06-09T16:30:25Z")

</div>

@nalimilan Thank You very much. I believe i was able to work out one the way to achieve the desired result

```julia
for c in unique(df[:Country])
    #df[Symbol(c)] = df[:Country] .== c
    df[Symbol(c)] = ones(Float64, size(df,1))
    for i in 1:size(df[:Country],1)
      if c == df[i,:Country]
        df[i,Symbol(c)] = 1.0
      else
        df[i,Symbol(c)] = 0.0
      end
    end
end

```

Kindly let me know if there is any better way to do it. Thank you.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [June 9, 2017, 5:00pm UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/10 "2017-06-09T17:00:30Z")

</div>

Just do this:

```julia
for c in unique(df[:Country])
    df[Symbol(c)] = UInt.(df[:Country] .== c)
end

```

or

```julia
for c in unique(df[:Country])
    df[Symbol(c)] = ifelse.(df[:Country] .== c, 1, 0)
end

```

The dot vectorized syntax ensures that no temporary vector will be created. But you can also keep the column as `Bool` as in many operations it will behave as expected: `false * 2 == 0`.

---

<div class="post-metadata">

**Author:** ![Saran\_S](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/saran_s/32/1155_2.png) [@Saran\_S](https://discourse.julialang.org/u/Saran_S)\
**Post date:** [June 9, 2017, 5:10pm UTC](https://discourse.julialang.org/t/dummy-encoding-one-hot-encoding-from-pooleddataarray/4167/11 "2017-06-09T17:10:13Z")

</div>

@nalimilan Thank You very very much. esp for the below tip.

> [@nalimilan](#):
>
> But you can also keep the column as Bool as in many operations it will behave as expected: false \* 2 == 0
