# Reading DataFrames from JLD2 files

**URL:** <https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543>\
**Category:** General Usage\
**Tags:** question, jld2, dataframes\
**Created:** [October 10, 2022, 11:25pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543 "2022-10-10T23:25:05Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![ehsani63](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ehsani63/32/38199_2.png) [@ehsani63](https://discourse.julialang.org/u/ehsani63)\
**Post date:** [October 10, 2022, 11:25pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/1 "2022-10-10T23:25:05Z")

</div>

The results of my simulations are saved as DataFrames in JLD2 files:

```julia
struct Sim
    params::Dict
    adf::DataFrame
    model::ABM
end
# Running simulations and generating results...
...

sim = Sim(d, adf, model)

# convert struct to dict
sim2d = struct2dict(sim)

# save the dict to name, using DrWatson safesave
safesave(fname, sim2d)

```

Until tonight, everything was working fine, but now I get the following warning when I try to load the data:

```julia-repl
julia> log_file = jldopen(fnamef)                                                                                                                                                       
...
├─🔢 params
├─🔢 adf
└─🔢 model

julia> adf = log_file["adf"]
┌ Warning: saved type DataFrame is missing field metadata in workspace type; reconstructing
└ @ JLD2 ~/.julia/packages/JLD2/HnW0g/src/data/reconstructing_datatypes.jl:164

```

`DataFrames` package is already loaded via the `using` command and `DataFrame`datatype is defined in the workspace. I am confused about the meaning of the warning, what could be the source of the problem and how to deal with it. I appreciate your comments and help.

I am working with `Julia@1.8.2`, `JLD2@0.4.25` and `DataFrames@1.4.1`.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [October 10, 2022, 11:29pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/2 "2022-10-10T23:29:14Z")

</div>

Perhaps it is this change:

```julia
DataFrame is now a mutable struct and has three new fields 
metadata, colmetadata, and allnotemetadata; 
this change makes DataFrame objects serialized under earlier 
versions of DataFrames.jl incompatible with 
version 1.4 (#3055)

```

For anyone needing to convert `DataFrame` objects from 1.3 to 1.4 version (e.g. if you serialized your objects for short-term storage).

The easiest solution is to use `Tables.columntable` on `DataFrame` objects created under DataFrames.jl 1.3 version. Then serialize them. Next upgrade DataFrames.jl to 1.4. Deserialize the `NamedTuple`, and next transform it back to `DataFrame`.

---

<div class="post-metadata">

**Author:** ![ehsani63](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ehsani63/32/38199_2.png) [@ehsani63](https://discourse.julialang.org/u/ehsani63)\
**Post date:** [October 10, 2022, 11:38pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/3 "2022-10-10T23:38:56Z")

</div>

Great! That was indeed the problem. Downgrading the `DataFrames` solved the issue. Now I need to convert all of the saved DataFrames to 1.4! 🥴

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [October 10, 2022, 11:41pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/4 "2022-10-10T23:41:25Z")

</div>

But there is no hurry. Just set the compat entry of DataFrames of your project to ~1.3 and make the data conversion whenever you have some spare time…

---

<div class="post-metadata">

**Author:** ![vladdez](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vladdez/32/45806_2.png) [@vladdez](https://discourse.julialang.org/u/vladdez)\
**Post date:** [January 12, 2023, 12:59pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/5 "2023-01-12T12:59:38Z")

</div>

Hi! Could you please elaborate more thoroughly what exactly should be done? What should be converted to what and how? And what and how should be serialized?

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [January 12, 2023, 1:04pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/6 "2023-01-12T13:04:06Z")

</div>

Perhaps convert the dataframes to Tables using the old dataframe package and serialize them as Tables?  
Then, using the new DataFrames package read the Tables from the jld2 file and convert them into the new DataFrames.

---

<div class="post-metadata">

**Author:** ![vladdez](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vladdez/32/45806_2.png) [@vladdez](https://discourse.julialang.org/u/vladdez)\
**Post date:** [January 12, 2023, 2:16pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/7 "2023-01-12T14:16:42Z")

</div>

![image](https://global.discourse-cdn.com/julialang/original/3X/c/e/ce3028c417a981088a53dc949ef64468165007e7.png)  
Like this? But there is no such format…

---

<div class="post-metadata">

**Author:** ![bkamins](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bkamins/32/208538_2.png) [@bkamins](https://discourse.julialang.org/u/bkamins)\
**Post date:** [January 12, 2023, 2:27pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/8 "2023-01-12T14:27:58Z")

</div>

on old DataFrames.jl:

1. `to_save = Tables.columntable(old_dataframe)`
2. serialize `to_save`.

Then on new DataFrames.jl:

1. deserialize to `from_save`
2. new\_dataframe = DataFrame(from\_save)`

---

<div class="post-metadata">

**Author:** ![JonasIsensee](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jonasisensee/32/4704_2.png) [@JonasIsensee](https://discourse.julialang.org/u/JonasIsensee)\
**Post date:** [January 13, 2023, 2:33pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/9 "2023-01-13T14:33:12Z")

</div>

Here’s a direct way to load “old” DataFrames with the new version.

Example:

1. create file with old DataFrames:

```julia
pkg> activate --temp
pkg> add DataFrames#v1.3.6
pkg> add JLD2 DelimitedFiles
julia> using JLD2, DelimitedFiles, DataFrames
julia> data, header = readdlm(joinpath(dirname(pathof(DataFrames)),
                                              "..", "docs", "src", "assets", "iris.csv"),
                                     ',', header=true);

julia> iris_raw = DataFrame(data, vec(header))
julia> iris = identity.(iris_raw)
lia> jldsave("old_dataframes.jld2"; iris)

```

Loading this in a new session with new DataFrames:

```julia
pkg> activate --temp
pkg> add JLD2 DataFrames
julia> using JLD2, DataFrames
julia> function Base.convert(::Type{DataFrame}, nt::NamedTuple)
           return DataFrame(nt.columns, nt.colindex)
       end

julia> load("old_dataframes.jld2", "iris"; 
    typemap = Dict("DataFrames.DataFrame" => JLD2.Upgrade(DataFrame))
    )

150×5 DataFrame
 Row │ SepalLength SepalWidth PetalLength PetalWidth Species        
     │ Float64 Float64 Float64 Float64 SubStrin…      
─────┼──────────────────────────────────────────────────────────────────
   1 │ 5.1 3.5 1.4 0.2 Iris-setosa
   2 │ 4.9 3.0 1.4 0.2 Iris-setosa
   3 │ 4.7 3.2 1.3 0.2 Iris-setosa
   4 │ 4.6 3.1 1.5 0.2 Iris-setosa
  ⋮ │ ⋮ ⋮ ⋮ ⋮ ⋮
 147 │ 6.3 2.5 5.0 1.9 Iris-virginica
 148 │ 6.5 3.0 5.2 2.0 Iris-virginica
 149 │ 6.2 3.4 5.4 2.3 Iris-virginica
 150 │ 5.9 3.0 5.1 1.8 Iris-virginica
                                                        142 rows omitted

```

(note that the constructor for the DataFrame I used here explicitly states that it shouldn’t be used directly…)

What’s happening here is that the `typemap` argument tells `JLD2` to watch out for stored structs  
with the name `"DataFrames.DataFrame"` and should use a special reconstruction method, i.e.  
load the fields as a named tuple and call `convert` to get a `DataFrame`.

---

<div class="post-metadata">

**Author:** ![vladdez](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vladdez/32/45806_2.png) [@vladdez](https://discourse.julialang.org/u/vladdez)\
**Post date:** [January 16, 2023, 12:58pm UTC](https://discourse.julialang.org/t/reading-dataframes-from-jld2-files/88543/10 "2023-01-16T12:58:00Z")

</div>

thanks, that works!

`to\_save1 = Tables.columntable(evt\_e)

file1 = File(format"JLD2", “data/evt\_e.jld2”)

save(file1, “1”, to\_save1)  
`
