# Best practice to unstack a dataframe with a lot of columnes?

**URL:** <https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851>\
**Category:** General Usage\
**Tags:** question, dataframes\
**Created:** [March 11, 2020, 8:26pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851 "2020-03-11T20:26:27Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 11, 2020, 8:26pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/1 "2020-03-11T20:26:28Z")

</div>

so, unlike normal ordering (like what Italy CDC did, JHU people [decided](https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_time_series/time_series_19-covid-Confirmed.csv) to put dates in column names.

```julia
#first(df_us, 3)

	State	Country	Lat	Long	1/22/20	1/23/20	1/24/20	1/25/20	1/26/20	1/27/20	1/28/20	1/29/20	1/30/20	1/31/20	2/1/20	2/2/20	2/3/20	2/4/20	2/5/20	2/6/20	2/7/20	2/8/20	2/9/20	2/10/20	2/11/20	2/12/20	2/13/20	2/14/20	2/15/20	2/16/20	2/17/20	2/18/20	2/19/20	2/20/20	2/21/20	2/22/20	2/23/20	2/24/20	2/25/20	2/26/20	2/27/20	2/28/20	2/29/20	3/1/20	3/2/20	3/3/20	3/4/20	3/5/20	3/6/20	3/7/20	3/8/20	3/9/20	3/10/20
	String⍰	String	Float64	Float64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64	Int64
1	Washington	US	47.4009	-121.49	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	267
2	New York	US	42.1657	-74.9481	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	173
3	California	US	36.1162	-119.682	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	0	144

```

now, the question is how to stuck unstack so that, in the end, I can do some sane plotting such as

```julia
@df df plot(:Dates, :Number, group=:State)

```

Currently I have to look by `eachrow` and prepare dates ahead by collecting `names(df)`, access data point by `collect(values(row[5:end]))`.

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 11, 2020, 8:29pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/2 "2020-03-11T20:29:13Z")

</div>

P.S, for Italy’s data:

```julia
df_italy = HTTP.get("https://raw.githubusercontent.com/pcm-dpc/COVID-19/master/dati-json/dpc-covid19-ita-province.json").body |>
JSON3.read |> DataFrame
@df df_italy plot(:data, :totale_casi, group=:sigla_provincia, xrotation=40)

```

produce [this](http://uaf-8.t2.ucsd.edu/~jling/dump/italy.html)

---

<div class="post-metadata">

**Author:** ![bashonubuntu](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bashonubuntu](https://discourse.julialang.org/u/bashonubuntu)\
**Post date:** [March 11, 2020, 8:31pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/3 "2020-03-11T20:31:16Z")

</div>

Maybe use `IndexedTables.stack()`?

[https://juliacomputing.github.io/JuliaDB.jl/latest/api/](https://juliacomputing.github.io/JuliaDB.jl/latest/api/)

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [March 11, 2020, 8:45pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/4 "2020-03-11T20:45:52Z")

</div>

Why not to use DataFrames functions?  
[Reshaping · DataFrames.jl](https://juliadata.github.io/DataFrames.jl/stable/man/reshaping_and_pivoting/#Reshaping-and-Pivoting-Data-1).

I’ve used this approach: [Covid-19 julia reshape · GitHub](https://gist.github.com/Arkoniak/a5c2cd23be353c54938fecf7317518bd)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [March 11, 2020, 8:52pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/5 "2020-03-11T20:52:25Z")

</div>

Yea I think you should be able to just use `stack` from `DataFrames` for this. Something like (guessing)

```julia
stack(df,5:end,:State,:Date,:Number)

```

---

<div class="post-metadata">

**Author:** ![jling](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jling/32/212909_2.png) [@jling](https://discourse.julialang.org/u/jling)\
**Post date:** [March 11, 2020, 9:01pm UTC](https://discourse.julialang.org/t/best-practice-to-unstack-a-dataframe-with-a-lot-of-columnes/35851/6 "2020-03-11T21:01:54Z")

</div>

> [@jling](#):
>
> how to stuck unstack

yes I was refering to Dataframes functions, your approach makes sense, thx
