# Creating columns in DataFrame via loops

**URL:** <https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287>\
**Category:** New to Julia\
**Tags:** question, dataframes\
**Created:** [June 27, 2025, 9:22pm UTC](https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287 "2025-06-27T21:22:34Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Snowy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/snowy/32/36765_2.png) [@Snowy](https://discourse.julialang.org/u/Snowy)\
**Post date:** [June 27, 2025, 9:22pm UTC](https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287/1 "2025-06-27T21:22:35Z")

</div>

Hi all,  
Stuck on what is probably a super simple problem. It’s the end of the week at 4pm on a Friday so my brain is just no longer working 😊

I want to create a columns on an existing data frame with a specified value.

Here is my code:

```julia
df = DataFrame("A"=>1:10)
category = ["cat1","cat2","cat3","cat4","cat5"]
cat_values = [0.203842327,0.210149485,0,0.070243409,0.034921919]

for i in category, x in cat_values
    df[:,i] .= cat_values[x]
end
#this produces error: ERROR: ArgumentError: invalid index: 0.203842327 of type Float64

```

I expected the behavior to do this (e.g. after a single loop):

```julia
df[:,:cat1] .= cat_values[1]
return(df)

```

Appreciate if anyone can point me in the right direction. It’s been a long day.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [June 27, 2025, 9:34pm UTC](https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287/2 "2025-06-27T21:34:47Z")

</div>

```julia
for (i,x) in zip(category, cat_values)
  df[:,i] = repeat([x], 10) # 10 values required to match column A
end

10×6 DataFrame
 Row │ A cat1 cat2 cat3 cat4 cat5
     │ Int64 Float64 Float64 Float64 Float64 Float64
─────┼──────────────────────────────────────────────────────────
   1 │ 1 0.203842 0.210149 0.0 0.0702434 0.0349219
   2 │ 2 0.203842 0.210149 0.0 0.0702434 0.0349219
   3 │ 3 0.203842 0.210149 0.0 0.0702434 0.0349219
   4 │ 4 0.203842 0.210149 0.0 0.0702434 0.0349219
   5 │ 5 0.203842 0.210149 0.0 0.0702434 0.0349219
   6 │ 6 0.203842 0.210149 0.0 0.0702434 0.0349219
   7 │ 7 0.203842 0.210149 0.0 0.0702434 0.0349219
   8 │ 8 0.203842 0.210149 0.0 0.0702434 0.0349219
   9 │ 9 0.203842 0.210149 0.0 0.0702434 0.0349219
  10 │ 10 0.203842 0.210149 0.0 0.0702434 0.0349219

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [June 27, 2025, 10:20pm UTC](https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287/3 "2025-06-27T22:20:15Z")

</div>

Another method (might be slower due to use of function closure):

```julia
foreach((n,v)->df[!,n] .= v, category, cat_values)

```

---

<div class="post-metadata">

**Author:** ![Benny](https://avatars.discourse-cdn.com/v4/letter/b/49beb7/32.png) [@Benny](https://discourse.julialang.org/u/Benny)\
**Post date:** [June 27, 2025, 10:51pm UTC](https://discourse.julialang.org/t/creating-columns-in-dataframe-via-loops/130287/4 "2025-06-27T22:51:40Z")

</div>

To walk through your code…

> [@Snowy](#):
>
> I expected the behavior to do this (e.g. after a single loop):
> 
> ```julia
> df[:,:cat1] .= cat_values[1]
> 
> ```

Let’s check that against your first iteration (that’s not enough debugging generally but it’s enough here). The first `i in category` is `"cat1"`, and the first `x in cat_values` is `0.203842327`, so the first iteration does:

```julia
julia> df[:, "cat1"] .= cat_values[0.203842327]
ERROR: ArgumentError: invalid index: 0.203842327 of type Float64

```

The same error, and it’s apparent we didn’t mean to index `cat_values` a 2nd time.

```julia
julia> df[:, "cat1"] .= 0.203842327
10-element Vector{Float64}:
 0.203842327
 0.203842327
 0.203842327
...

```

So let’s make that change in your loop and check the resulting `df`:

```julia
julia> for i in category, x in cat_values
         df[:,i] .= x
       end

julia> df
10×6 DataFrame
 Row │ A cat1 cat2 cat3 cat4 cat5
     │ Int64 Float64 Float64 Float64 Float64 Float64
─────┼──────────────────────────────────────────────────────────────
   1 │ 1 0.0349219 0.0349219 0.0349219 0.0349219 0.0349219
   2 │ 2 0.0349219 0.0349219 0.0349219 0.0349219 0.0349219
   3 │ 3 0.0349219 0.0349219 0.0349219 0.0349219 0.0349219
...

```

Well that’s not what we want either. We’re iterating through `category` columns and `cat_values` values, so what’s the problem? Let’s reference the loop docs:

> Multiple nested `for` loops can be combined into a single outer loop, forming the cartesian product of its iterables:
> 
> ```julia
> julia> for i = 1:2, j = 3:4
> println((i, j))
> end
> (1, 3)
> (1, 4)
> (2, 3)
> (2, 4)
> 
> ```

So instead of 5 iterations of `category` and `cat_values` in parallel, we had 5x5=25 iterations of the Cartesian product of their elements. The order made it so that we filled each column with successive values of `cat_values` until the last `0.0349219`. To iterate 2 or more sequences in parallel (until the shortest one is exhausted), we can use `zip`:

```julia
julia> for (i,x) in zip(category, cat_values)
           df[:,i] .= x
       end

julia> df
10×6 DataFrame
 Row │ A cat1 cat2 cat3 cat4 cat5
     │ Int64 Float64 Float64 Float64 Float64 Float64
─────┼──────────────────────────────────────────────────────────
   1 │ 1 0.203842 0.210149 0.0 0.0702434 0.0349219
   2 │ 2 0.203842 0.210149 0.0 0.0702434 0.0349219
   3 │ 3 0.203842 0.210149 0.0 0.0702434 0.0349219
...

```

Seems right.
