# Writing to DataFrames and CSV files

**URL:** <https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238>\
**Category:** General Usage\
**Tags:** dataframes, csv\
**Created:** [March 20, 2020, 12:55am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238 "2020-03-20T00:55:37Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 12:55am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/1 "2020-03-20T00:55:37Z")

</div>

Hi, I am looking to write output of a function for different runs to one CSV file. Using the DelimitedFiles package, I can write results to separate CSV files for each run. However, I want to write results from all runs to one CSV file either in separate rows or separate columns.

As an example, I have given some code below. This will create separate CSV file for each run.

To write all results to one CSV file I use `CSV.write("TestCSVPack.csv", DataFrame(result))` but I get the following error message  
**ArgumentError: ‘Base.Generator{UnitRange{Int64},var"#13#14"{Array{Int64,1}}}’ iterates ‘Int64’ values, which doesn’t satisfy the Tables.jl `AbstractRow` interface**

```julia
using CSV, DataFrames, DelimitedFiles
function Addition(x,y,i)
    return x[:,i]+y[:,i]
end
x = [1 2 3; 6 8 9]
y = [1 3 1; 10 12 5]
for i = 1:size(x,2)
    result = Addition(x,y,i)
# use DelimitedFiles to create a separate CSV file for each run
    writedlm( "TestDelimited$i.csv", result)
# use CSV package to create a CSV file for all results. The following line will give an error
   CSV.write("TestCSVPack.csv", DataFrame(result))
end

```

**Is there a simple way to fix it, either using the CSV package or another package capable of writing CSV files?**

---

<div class="post-metadata">

**Author:** ![alejandromerchan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/alejandromerchan/32/10500_2.png) [@alejandromerchan](https://discourse.julialang.org/u/alejandromerchan)\
**Post date:** [March 20, 2020, 2:07am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/2 "2020-03-20T02:07:42Z")

</div>

Maybe not exactly what you want, but this works

```julia
open("myfile.txt", "w") do io
    for i = 1:size(x,2)
        result = Addition(x,y,i)
        write(io, string(result))
    end
end

```

Played some with CSV and DataFrames, but couldn’t make it work thus far.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 20, 2020, 2:27am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/3 "2020-03-20T02:27:45Z")

</div>

one thing you can do is `vcat` all the results together before writing

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [March 20, 2020, 6:21am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/4 "2020-03-20T06:21:40Z")

</div>

I would concatenate the results into a `DataFrame`, also keeping track of the `i` indexes for later use. Something like

```julia
julia> mapreduce(i -> (result = addition(x, y, i);
       DataFrame(i = fill(i, length(result)), result = result)),
       vcat, 1:size(x, 2))
6×2 DataFrame
│ Row │ i │ result │
│ │ Int64 │ Int64 │
├─────┼───────┼────────┤
│ 1 │ 1 │ 2 │
│ 2 │ 1 │ 16 │
│ 3 │ 2 │ 5 │
│ 4 │ 2 │ 20 │
│ 5 │ 3 │ 4 │
│ 6 │ 3 │ 14 │

```

which you can then save using `CSV.write`.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 20, 2020, 8:05am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/5 "2020-03-20T08:05:44Z")

</div>

There are two different problems in your approach.  
First one is the error message you receive, and you see it because there is no `DataFrame` constructor of the column, since `typeof(result)` is `Vector{Int64} `. You can avoid it by either giving explicit name to this column, or by turning it into `Matrix` transposing it. For example

```julia
result = Addition(x, y, 1)

# Option 1

julia> DataFrame(:x => result)
2×1 DataFrame
│ Row │ x │
│ │ Int64 │
├─────┼───────┤
│ 1 │ 2 │
│ 2 │ 16 │

# Option 2
julia> DataFrame(result')
1×2 DataFrame
│ Row │ x1 │ x2 │
│ │ Int64 │ Int64 │
├─────┼───────┼───────┤
│ 1 │ 2 │ 16 │

```

Second problem is writing to the same file many times, it can be solved by using `append` flag.

```julia
using CSV
using DataFrames

function Addition(x,y,i)
    return x[:,i]+y[:,i]
end
x = [1 2 3; 6 8 9]
y = [1 3 1; 10 12 5]
for i = 1:size(x,2)
    result = Addition(x,y,i)
    CSV.write("TestCSVPack.csv", DataFrame(:x => result), append = i != 1, writeheader = false)
end

```

with content of `TestCSVPack.csv`

```nohighlight
sh> cat TestCSVPack.csv
2
16
5
20
4
14

```

As additional note, if your data is homogeneous as in your example, you may omit DataFrames since they introduce some overhead and use other `Tables.jl` sinks, for example (here I also write row by row)

```julia
using CSV
using Tables

for i = 1:size(x,2)
    result = Addition(x,y,i)
   CSV.write("TestCSVPack.csv", Tables.table(result'), append = i != 1, writeheader = false)
end

```

with contents

```nohighlight
sh> cat TestCSVPack.csv
2,16
5,20
4,14

```

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [March 20, 2020, 8:19am UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/6 "2020-03-20T08:19:06Z")

</div>

Showing my ignorance here… would it be worth considering other file formats?  
for instance HDF5 [GitHub - JuliaIO/HDF5.jl: Save and load data in the HDF5 file format from Julia](https://github.com/JuliaIO/HDF5.jl)

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 12:35pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/7 "2020-03-20T12:35:27Z")

</div>

Thanks. This works fine for the example I gave. I am not sure if this approach requires storing results from all runs in to an array before writing to a dataframe. So it may not work significantly large sized problems.

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 12:51pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/8 "2020-03-20T12:51:33Z")

</div>

Thanks. Your suggestion is perfect for my needs as I think it does not require holding results from all the runs in the memory at the same time. **Can I check how can I record i (run number) in to the CSV as in the response given by Tamas?**

For e.g. using the code below records i but I cannot work out how it can be used based on your suggestion

```julia
df=mapreduce(i -> (result = Addition(x, y, i); DataFrame(I= fill(i, length(result)), Result = result)),vcat, 1:size(x, 2))
CSV.write("TestCSVPack.csv", df)

```

This produces the following DataFrame

```julia
6×2 DataFrame
│ Row │ I │ Result │
│ │ Int64 │ Int64 │
├─────┼───────┼────────┤
│ 1 │ 1 │ 2 │
│ 2 │ 1 │ 16 │
│ 3 │ 2 │ 5 │
│ 4 │ 2 │ 20 │
│ 5 │ 3 │ 4 │
│ 6 │ 3 │ 14 │

```

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 1:03pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/9 "2020-03-20T13:03:26Z")

</div>

Certainly it is worth considering other file formats. I am not familiar with HDF5 format but it looks interesting. I cannot find much documentation with this package so it may take me sometime to work out how to write results for the example given in the question.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 20, 2020, 2:26pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/10 "2020-03-20T14:26:26Z")

</div>

You should proceed in the same way as was proposed by Tamas, just add new column in resulting DataFrame

```julia
for i = 1:size(x,2)
    result = Addition(x,y,i)
    CSV.write("TestCSVPack.csv", DataFrame(:n => i, :x => result), append = i != 1, writeheader = false)
end

```

N.B.: as you can see, I’ve used just `i`, not `fill(i, length(result))` because `DataFrame` constructor auto cycle this column, but both versions are valid. Also, you can remove `writeheader` argument if you want to, because `CSV` does not create additional header rows in `append` mode.

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 3:40pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/11 "2020-03-20T15:40:27Z")

</div>

Thanks.  
Lastly, how can I add headings to the CSV /DataFrame i.e. to have “i” and “Result” as headings for the data (to appear once at the top)?

Using this code would write headings for each run, but I want it to appear once  
`CSV.write("TestCSVPack.csv", DataFrame(I = i, Result = result), append = i != 1, writeheader = true)`

```julia
|I|Result|
|1|2|
|1|16|
|I|Result|
|2|5|
|2|20|
|I|Result|
|3|4|
|3|14|

```

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 20, 2020, 4:40pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/12 "2020-03-20T16:40:31Z")

</div>

Just remove `writeheader` completely.

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 5:21pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/13 "2020-03-20T17:21:01Z")

</div>

Thank you

---

<div class="post-metadata">

**Author:** ![Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/smith/32/13567_2.png) [@Smith](https://discourse.julialang.org/u/Smith)\
**Post date:** [March 20, 2020, 9:16pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/14 "2020-03-20T21:16:45Z")

</div>

Thanks for looking in to jt

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [August 24, 2020, 12:38pm UTC](https://discourse.julialang.org/t/writing-to-dataframes-and-csv-files/36238/15 "2020-08-24T12:38:31Z")

</div>

Sorry for re-opening this old thread, but it seems that HDF5.jl does not directly work together with DataFrames. The package HDFTables.jl, which seems to be such an interface, is not maintained anymore, as far as I can see.  
Is there a simple way to open HDF files (saved in Pandas) in Julia DataFrames?
