# Convert collection (Array, DataFrame, ...) to concrete eltype

**URL:** <https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168>\
**Category:** New to Julia\
**Created:** [August 29, 2019, 2:53pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168 "2019-08-29T14:53:14Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [August 29, 2019, 2:53pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/1 "2019-08-29T14:53:15Z")

</div>

Suppose I have a collection, e.g. DataFrame, with `Any` eltype but all elements having same concrete type:

```julia
df = DataFrame(a=Any[1, 2, 3])

```

For further processing I need to make it type-stable, but don’t see how to do that. Any obvious way I’m missing here?

Such situation occurs when reading a “dirty” dataset with all kind of wrong values, and cleaning it afterwards.

---

<div class="post-metadata">

**Author:** ![kevbonham](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kevbonham/32/216165_2.png) [@kevbonham](https://discourse.julialang.org/u/kevbonham)\
**Post date:** [August 29, 2019, 3:12pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/2 "2019-08-29T15:12:50Z")

</div>

Perhaps not the most efficient, but:

```julia
julia> df = DataFrame(a=Any[1, 2, 3], b=Any[1., 2, 3])
3×2 DataFrame
│ Row │ a │ b │
│ │ Any │ Any │
├─────┼─────┼─────┤
│ 1 │ 1 │ 1.0 │
│ 2 │ 2 │ 2 │
│ 3 │ 3 │ 3 │

julia> for n in names(df)
           df[!,n] = [x for x in df[!,n]]
       end

julia> df
3×2 DataFrame
│ Row │ a │ b │
│ │ Int64 │ Real │
├─────┼───────┼──────┤
│ 1 │ 1 │ 1.0 │
│ 2 │ 2 │ 2 │
│ 3 │ 3 │ 3 │

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [August 29, 2019, 3:29pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/3 "2019-08-29T15:29:31Z")

</div>

I might misunderstand but will this do?

```julia
julia> using DataFrames

julia> df = DataFrame(a=Any[1, 2, 3])

julia> df.a = Int64.(df.a)

julia> df
3×1 DataFrame
│ Row │ a │
│ │ Int64 │
├─────┼───────┤
│ 1 │ 1 │
│ 2 │ 2 │
│ 3 │ 3 │

```

or maybe

```julia
eltype(df.a[1]).(df.a)

```

if you want it to be more generic (and can rely on that first value…)

---

<div class="post-metadata">

**Author:** ![aaowens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaowens/32/12101_2.png) [@aaowens](https://discourse.julialang.org/u/aaowens)\
**Post date:** [August 29, 2019, 6:56pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/4 "2019-08-29T18:56:12Z")

</div>

I find comprehensions tend to solve this automatically for me

```julia
julia> using DataFrames

julia> df = DataFrame(a=Any[1, 2, 3])
3×1 DataFrame
│ Row │ a │
│ │ Any │
├─────┼─────┤
│ 1 │ 1 │
│ 2 │ 2 │
│ 3 │ 3 │
julia> [aa for aa in df.a]
3-element Array{Int64,1}:
 1
 2
 3

```

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [August 29, 2019, 7:19pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/5 "2019-08-29T19:19:35Z")

</div>

Thanks for suggestions! For now comprehensions seems like the best easy choice

```julia
for n in names(df)
    df[!,n] = [x for x in df[!,n]]
end

```

Explicitly using type of the first element like `typeof(df.a[1]).(df.a)` (note `typeof` instead of `eltype` as was suggested - so that it works for arrays as well) is definitely less general. E.g. it doesn’t work for `Union{..., Nothing}` which is pretty common, and other small unions which are handled well by comprehensions.

For larger datasets where performance is important it would be better to have a helper function to skip columns which already have proper types. Unfortunately, I don’t think it’s possible to determine if the type is correct without checking all values anyway…

---

<div class="post-metadata">

**Author:** ![aaowens](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aaowens/32/12101_2.png) [@aaowens](https://discourse.julialang.org/u/aaowens)\
**Post date:** [August 29, 2019, 8:30pm UTC](https://discourse.julialang.org/t/convert-collection-array-dataframe-to-concrete-eltype/28168/6 "2019-08-29T20:30:33Z")

</div>

I wonder if this would be a nice feature to be built into DataFrames. Something like `narrowtypes!(df)` which in simplest form does your loop, but could be made more efficient by skipping any column which already has a concrete type. Like this,

```julia
for n in names(df)
    isconcretetype(eltype(df[!, n])) && continue
    df[!,n] = [x for x in df[!,n]]
end

```
