# DataFrames: ByRow fails in transform with PooledArrays after CSV.read

**URL:** <https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957>\
**Category:** Data\
**Tags:** question\
**Created:** [August 6, 2021, 5:26pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957 "2021-08-06T17:26:02Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 6, 2021, 5:26pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/1 "2021-08-06T17:26:02Z")

</div>

I am having an issue with a `transform` operation on a `DataFrame` that is read from a .csv using `CSV`. I don’t have the following issue when my dataframe has few rows because the importing from the .csv file uses `SentinelArrays`. However, when I have a larger dataframe, `PooledArrays` are used.

The problem is the following: I have a column with `missing` values. I would like to replace each `missing` value with a unique symbol using `gensym()`. The transformation works as expected on the freshly create dataframe, but when I save it as a .csv and read it into a new DataFrame, all `missing` values are replaced with the same `gensym()`. Any thoughts here? If I try to use `map` instead of `transform`, the same issue occurs.

```julia
julia> using CSV, DataFrames

julia> x=DataFrame(A=vcat("C",repeat([missing],1000)),B=3)
1001×2 DataFrame
  Row │ A B     
      │ String? Int64 
──────┼────────────────
    1 │ C 3
    2 │ missing 3
    3 │ missing 3
    4 │ missing 3
    5 │ missing 3
  ⋮ │ ⋮ ⋮
  995 │ missing 3
  996 │ missing 3
  997 │ missing 3
  998 │ missing 3
  999 │ missing 3
 1000 │ missing 3
 1001 │ missing 3
       955 rows omitted

julia> CSV.write("x.csv",x)
"x.csv"

julia> transform(x,:A => ByRow(i -> ismissing(i) ? gensym() : i) => :A)
1001×2 DataFrame
  Row │ A B     
      │ Any Int64 
──────┼───────────────
    1 │ C 3
    2 │ ##1754 3
    3 │ ##1755 3
    4 │ ##1756 3
    5 │ ##1757 3
     ⋮ │ ⋮ ⋮
  995 │ ##2747 3
  996 │ ##2748 3
  997 │ ##2749 3
  998 │ ##2750 3
  999 │ ##2751 3
 1000 │ ##2752 3
 1001 │ ##2753 3
      955 rows omitted

julia> y=CSV.read("x.csv",DataFrame);

julia> transform(y,:A => ByRow(i -> ismissing(i) ? gensym() : i) => :A)
1001×2 DataFrame
  Row │ A B     
      │ Any Int64 
──────┼───────────────
    1 │ C 3
    2 │ ##2755 3
    3 │ ##2755 3
    4 │ ##2755 3
    5 │ ##2755 3
  ⋮ │ ⋮ ⋮
  995 │ ##2755 3
  996 │ ##2755 3
  997 │ ##2755 3
  998 │ ##2755 3
  999 │ ##2755 3
 1000 │ ##2755 3
 1001 │ ##2755 3
      955 rows omitted

```

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 6, 2021, 5:35pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/2 "2021-08-06T17:35:23Z")

</div>

So, I found I can use the keyword argument `pool = false` when reading the .csv to fix this issue.

@bkamins, should I submit an issue on DataFrames.jl for this? or is this behavior expected when applying `transform` to a `string?` column that is pooled?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 6, 2021, 5:48pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/3 "2021-08-06T17:48:48Z")

</div>

Can you try to make an MWE that omits DataFrames and file it at CSV.jl?

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 6, 2021, 7:25pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/4 "2021-08-06T19:25:25Z")

</div>

So it’s a CSV.jl issue? I thought it would be a DataFrames.jl issue since the problem occurs when trying to transform the data because it has been compressed by CSV.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [August 6, 2021, 7:27pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/5 "2021-08-06T19:27:36Z")

</div>

It’s a SentinalArrays issue, and therefore a CSV issue I think. Note that `ByRow(f)(x)` is just going to do `f.(x)`, so `ByRow` isn’t doing anything unique with this.

---

<div class="post-metadata">

**Author:** ![hdavid16](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hdavid16/32/11531_2.png) [@hdavid16](https://discourse.julialang.org/u/hdavid16)\
**Post date:** [August 6, 2021, 7:57pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/6 "2021-08-06T19:57:42Z")

</div>

It actually has to do with PooledArrays. Since these are compressed, I’m guessing all `missing` items are compressed to the same object, which is why they get replaced with the same `Symbol` with `gensym()`.

So maybe `pool=false` should be the default, but perhaps there are other (better) reasons why `pool=true` is the default.

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [August 6, 2021, 7:57pm UTC](https://discourse.julialang.org/t/dataframes-byrow-fails-in-transform-with-pooledarrays-after-csv-read/65957/7 "2021-08-06T19:57:54Z")

</div>

This is a DataFrames.jl/PooledArrays.jl issue. It is tracked in [https://github.com/JuliaData/PooledArrays.jl/issues/63](https://github.com/JuliaData/PooledArrays.jl/issues/63). I have opened [https://github.com/JuliaData/DataFrames.jl/issues/2834](https://github.com/JuliaData/DataFrames.jl/issues/2834) to make sure it is resolved soon in DataFrames.jl.

The reason is [https://github.com/JuliaData/PooledArrays.jl/blob/35ecfd186c5e0f1aba1fc278e93766f3258f9cc3/src/PooledArrays.jl#L307](https://github.com/JuliaData/PooledArrays.jl/blob/35ecfd186c5e0f1aba1fc278e93766f3258f9cc3/src/PooledArrays.jl#L307)
