# Is there a better way to do this? many calculated columns

**URL:** <https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111>\
**Category:** General Usage\
**Created:** [September 23, 2020, 5:33am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111 "2020-09-23T05:33:16Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [September 23, 2020, 5:33am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/1 "2020-09-23T05:33:16Z")

</div>

Creating a DataFrame with a large number of calculated columns.  
To make the logic as clear as possible, avoiding the vectorised operators and instead putting all logic within an eachrow loop.

```julia
new_DF = DataFrame()

for r = eachrow(base_DF)
    d = DotMap()

    d.Col1 = somelogic
    d.Col2 = somelogic
     ...
    d.Col20 = somelogic

    append!( new_DF, d ) 
end

```

I’d like to avoid the procedural elements:

```julia
        new_DF = DataFrame() `and`
        append!(new_DF, d)`

```

For a single vector there is the python style list comprehension [x for x in list]

Is there an equivalent for DataFrames ?

Can you create all the fields in a loop like this and they automatically output a DataFrame ?

---

<div class="post-metadata">

**Author:** ![remi-garcia](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/remi-garcia/32/7125_2.png) [@remi-garcia](https://discourse.julialang.org/u/remi-garcia)\
**Post date:** [September 23, 2020, 6:21am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/2 "2020-09-23T06:21:44Z")

</div>

Do you mean something like:

```nohighlight
julia> DataFrame([[i*2 for i in 12:20] [i+1 for i in 12:20]], ["name_1", "name_2"])
9×2 DataFrame
│ Row │ name_1 │ name_2 │
│ │ Int64 │ Int64 │
├─────┼────────┼────────┤
│ 1 │ 24 │ 13 │
│ 2 │ 26 │ 14 │
│ 3 │ 28 │ 15 │
│ 4 │ 30 │ 16 │
│ 5 │ 32 │ 17 │
│ 6 │ 34 │ 18 │
│ 7 │ 36 │ 19 │
│ 8 │ 38 │ 20 │
│ 9 │ 40 │ 21 │

```

[A B] corresponds to `hcat`. You could write `hcat([i*2 for i in 12:20], [i+1 for i in 12:20])` instead.

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [September 23, 2020, 6:34am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/3 "2020-09-23T06:34:49Z")

</div>

More like thes:

```julia
DataFrame(
       col_1 = 1:100,
       col_2 = col_1^2,
       col_3 = col_2 + 10
)

```

can you just define all the fields, with dependence on eachother

---

<div class="post-metadata">

**Author:** ![remi-garcia](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/remi-garcia/32/7125_2.png) [@remi-garcia](https://discourse.julialang.org/u/remi-garcia)\
**Post date:** [September 23, 2020, 7:40am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/4 "2020-09-23T07:40:45Z")

</div>

I don’t think that you can make a dependence on fields in one line.

That become harder and harder to read but:

```nohighlight
julia> DataFrame([collect(1:10) collect(1:10).^2 (collect(1:10).^2).+10])
10×3 DataFrame
│ Row │ x1 │ x2 │ x3 │
│ │ Int64 │ Int64 │ Int64 │
├─────┼───────┼───────┼───────┤
│ 1 │ 1 │ 1 │ 11 │
│ 2 │ 2 │ 4 │ 14 │
│ 3 │ 3 │ 9 │ 19 │
│ 4 │ 4 │ 16 │ 26 │
│ 5 │ 5 │ 25 │ 35 │
│ 6 │ 6 │ 36 │ 46 │
│ 7 │ 7 │ 49 │ 59 │
│ 8 │ 8 │ 64 │ 74 │
│ 9 │ 9 │ 81 │ 91 │
│ 10 │ 10 │ 100 │ 110 │

```

works. Now you have to hope that the compiler optimize the computations.

Maybe this could be a solution:

```nohighlight
julia> col_1 = 1:10; col_2 = col_1.^2; col_3 = col_2 .+ 10;

julia> DataFrame([col_1, col_2, col_3])
10×3 DataFrame
│ Row │ x1 │ x2 │ x3 │
│ │ Int64 │ Int64 │ Int64 │
├─────┼───────┼───────┼───────┤
│ 1 │ 1 │ 1 │ 11 │
│ 2 │ 2 │ 4 │ 14 │
│ 3 │ 3 │ 9 │ 19 │
│ 4 │ 4 │ 16 │ 26 │
│ 5 │ 5 │ 25 │ 35 │
│ 6 │ 6 │ 36 │ 46 │
│ 7 │ 7 │ 49 │ 59 │
│ 8 │ 8 │ 64 │ 74 │
│ 9 │ 9 │ 81 │ 91 │
│ 10 │ 10 │ 100 │ 110 │

julia> col_1 = col_2 = col_3 = nothing;

```

---

<div class="post-metadata">

**Author:** ![mariok90](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mariok90/32/7076_2.png) [@mariok90](https://discourse.julialang.org/u/mariok90)\
**Post date:** [September 23, 2020, 8:11am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/5 "2020-09-23T08:11:03Z")

</div>

You can also construct DataFrames with NamedTuples:

```julia
DataFrame(
       (col_1 = i,
       col_2 = i^2,
       col_3 = i^2 + 10) for i in 1:100
)

```

For more complex operations you could define functions beforehand.

```julia
func1(x) = x^2
func2(x) = func1(x) + 10

DataFrame(
       (col_1 = i,
       col_2 = func1(i),
       col_3 =func2(i)) for i in 1:100
)

```

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [September 23, 2020, 11:55am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/6 "2020-09-23T11:55:57Z")

</div>

Depending on exactly what the operations do, you could use `select(base_DF, ...)` or `combine(base_DF, ...)`.

---

<div class="post-metadata">

**Author:** ![Lincoln\_Hannah](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lincoln_hannah/32/19198_2.png) [@Lincoln\_Hannah](https://discourse.julialang.org/u/Lincoln_Hannah)\
**Post date:** [September 24, 2020, 12:40am UTC](https://discourse.julialang.org/t/is-there-a-better-way-to-do-this-many-calculated-columns/47111/7 "2020-09-24T00:40:27Z")

</div>

Thanks for you help everyone.  
I think I like this approach best: Vectorise an anonymous function that outputs a DotMap, then  
DataFrame the result. This allows each new column to use previous columns as input.  
It would be nice if you could just add new columns to the DataFrameRow variable but I don’t think this is possible.

```julia
new_DF = DataFrame( ( function( row ); d = DotMap()

    d.Col_1 = somelogic( row )
    d.Col_2 = somelogic( row, d.Col_1 )
     ...
    d.Col_20 = somelogic( row, d.Col_1, ... d.Col_19 )

    d end).(eachrow(base_DF)))

```
