# question about creating new columns in data frame from existing columns, 

**URL:** <https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [June 20, 2018, 5:56am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807 "2018-06-20T05:56:27Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![sarah-ji](https://avatars.discourse-cdn.com/v4/letter/s/e47774/32.png) [@sarah-ji](https://discourse.julialang.org/u/sarah-ji)\
**Post date:** [June 20, 2018, 5:56am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807/1 "2018-06-20T05:56:27Z")

</div>

Hi, I’m a little confused on what to do in this coding situation and would appreciate some help!

I’m trying to create new columns in a dataframe (D) as a linear combination of the existing columns in that dataframe, given a vector of coefficients (v).

Basically: If I have a vector of Float64 numbers, v\_i, and a 20 column data frame, D, of float64 values, how can I concatenate columns to the data frame D, where I am doing an element wise multiplication:  
D[:New] = v\_1_D\_1 + v\_2_D\_2 + v\_3_D\_3 + … v\_20_D\_20 ; where D\_i = ith column of D

So far I have started to write the function:

for i in 1:length(v)  
D[:New] += v[i]\*(D[:column\_names\_of\_D[i]])  
end

But I am unsure how to call the D[:column\_names\_of\_D[i]]) part…

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 20, 2018, 6:51am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807/2 "2018-06-20T06:51:49Z")

</div>

I would collect them as a matrix and then use `matrix * vector` multiplication, eg

```julia
using DataFrames

"""
Return columns that start with `prefix` as a matrix.
"""
function cols2matrix(df::DataFrame, prefix)
    matching_names = sort(filter(name -> startswith(String(name), String(prefix)),
                                 names(df)))
    hcat(getindex.(df, matching_names)...)
end

# make a dataframe
df = DataFrame(v_1 = randn(10), v_2 = randn(10), v_3 = randn(10))

M = cols2matrix(df, "v_")

M * [1, 2, 3]

```

But I would suggest that storing this kind of data in a matrix may be best in the first place.

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [June 20, 2018, 11:43am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807/3 "2018-06-20T11:43:25Z")

</div>

Just a note about the `hcat` command. `Matrix()` works just find on DataFrames and is probably more idiomatic.

I also am gonna work on a package that overloads some useful string commands, or writes wrappers for them, for common string operations that would be used on the names of dataframes columns.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [June 20, 2018, 11:49am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807/4 "2018-06-20T11:49:17Z")

</div>

> [@pdeffebach](#):
>
> `Matrix()` works just find on DataFrames

Excellent idea, as it allows more modular code:

```julia
using DataFrames

colnames_with_prefix(df::DataFrame, prefix) =
    sort(filter(name -> startswith(String(name), String(prefix)), names(df)))

df = DataFrame(v_1 = randn(10), v_2 = randn(10), v_3 = randn(10),
               a = rand(1:20, 10))

Matrix(df[colnames_with_prefix(df, :v_)]) * [1, 2, 3]

```

---

<div class="post-metadata">

**Author:** ![sarah-ji](https://avatars.discourse-cdn.com/v4/letter/s/e47774/32.png) [@sarah-ji](https://discourse.julialang.org/u/sarah-ji)\
**Post date:** [June 26, 2018, 4:16am UTC](https://discourse.julialang.org/t/question-about-creating-new-columns-in-data-frame-from-existing-columns/11807/5 "2018-06-26T04:16:34Z")

</div>

Thanks Tamas! 🙂
