# Most popular tabular/multidimensional data types in Julia

**URL:** <https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766>\
**Category:** New to Julia\
**Tags:** data, type, dataframes\
**Created:** [December 8, 2021, 12:12pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766 "2021-12-08T12:12:58Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 12:12pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/1 "2021-12-08T12:12:58Z")

</div>

I’m writing a Julia connector to a weather data API, and want some insight into the types of data preferred by the Julia community. It seems to me that the built-in DataFrames are very limited; furthermore, users might request data which would be better stored in multi-dimensional structures. I come from a Python background, so am looking to emulate things like Pandas and Xarray, but a) would like to avoid using PyCall explicitly and b) feel that leaning too heavily on frontends like Pandas.jl may not be the most Julian way of proceeding.

Any comments on your favourite packages and types for time-series or grids of one or more variables much appreciated! Additional info: the API responds with .csvs, which can be converted to built-in DataFrames as an intermediate step: I may well implement some kind of method/type switching options for the user, so ideally whatever packages you suggest would be able to take a native DataFrame as input.

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 8, 2021, 12:16pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/2 "2021-12-08T12:16:37Z")

</div>

> [@Eagertom](#):
>
> It seems to me that the built-in DataFrames are very limited;

in what way ?

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [December 8, 2021, 12:16pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/3 "2021-12-08T12:16:50Z")

</div>

> [@Eagertom](#):
>
> It seems to me that the built-in DataFrames are very limited;

Can you expand on this? It’s not clear from this post exactly what functionality you need and where DataFrames.jl falls short.

> [@Eagertom](#):
>
> users might request data which would be better stored in multi-dimensional structures.

It’s true DataFrames requires columns to be vectors, but so does pandas, correct? This Split-Apply-Combine strategy in DataFrames is very powerful and can emulate multi-dimensionality very well.

Maybe you want NamedArrays.jl Or AxisArrays.jl?

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [December 8, 2021, 12:27pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/4 "2021-12-08T12:27:03Z")

</div>

> [@Eagertom](#):
>
> It seems to me that the built-in DataFrames are very limited;

As others already mentioned above, it would be nice to understand your expectations from a data frame. DataFrames.jl is super powerful and does things very well when you have 2 dimensions (rows and columns). If you need more than 2 dimensions, consider other established data structures and formats such as HDF5, NetCDF, etc. We have all that in Julia as well.

If you want to integrate with the data science ecosystem in Julia, I would even try to create a custom data type that implements the Tables.jl interface. It would automatically work with statistical packages for example.

Notice that I have done a similar job connecting meteorological data here:

> **[GitHub - JuliaClimate/INMET.jl: Julia API to access data from the Instituto...](https://github.com/JuliaClimate/INMET.jl)**
>
> Julia API to access data from the Instituto Nacional de Metereologia (INMET) - GitHub - JuliaClimate/INMET.jl: Julia API to access data from the Instituto Nacional de Metereologia (INMET)

DataFrames.jl was the perfect fit.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [December 8, 2021, 12:28pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/5 "2021-12-08T12:28:00Z")

</div>

Another package that we wrote connecting meteorological data APIs:

[https://github.com/JuliaClimate/CDSAPI.jl](https://github.com/JuliaClimate/CDSAPI.jl)

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [December 8, 2021, 12:33pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/6 "2021-12-08T12:33:57Z")

</div>

> [@pdeffebach](#):
>
> It’s true DataFrames requires columns to be vectors, but so does pandas, correct? This Split-Apply-Combine strategy in DataFrames is very powerful and can emulate multi-dimensionality very well.

I think the OP is familiar with the multidimensional functionality of Pandas, which I personally find overly complicated: [MultiIndex / advanced indexing — pandas 2.1.3 documentation](https://pandas.pydata.org/pandas-docs/stable/user_guide/advanced.html)

They extend the concept of tables to hierarchical tables like we see in spreadsheets, where main columns are subdivided into other columns, …

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [December 8, 2021, 12:37pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/7 "2021-12-08T12:37:20Z")

</div>

I am working quite a lot with Pandas, but MultiIndex is a complexity nightmare without a clear unique use case.  
Glad that DataFrames.jl does not have it (and probably nobody misses it).

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [December 8, 2021, 12:39pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/8 "2021-12-08T12:39:37Z")

</div>

I second that. Maybe there are use cases out there, but I often find other data structures that are more appropriate and more intuitive when it comes to more than 2 dimensions. Particularly in Julia where Arrays are multidimensional by default and where so many other data structures are available implementing the `AbstractArray` interface.

---

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 12:57pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/9 "2021-12-08T12:57:57Z")

</div>

I actually agree with @juliohm that the multi-index functionality of Python’s Pandas is unnecessarily complicated. The principle thing I dislike about Julia DataFrames is that, as far as I can tell, I can’t set an index for my data manually, and that the index resets each time a new view is created. Finding that people disagree with me is valuable in its own right: if the community at large uses DataFrames heavily then I ought to implement a DataFrames solution!

@pdeffebach by multi-dimensional I mean something like netCDF data, where a 3D grid might be required to either store 3D data or to store 2D data for a range of variables. A 4th dimension is often also applicable to 3D grid time-series. Xarray is a nice package for this in Python, so I’m wondering what the Julia community uses.

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [December 8, 2021, 1:07pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/10 "2021-12-08T13:07:16Z")

</div>

The question is what is the use case for an index in Pandas? It is essentially an additional column with a special behavior in some functions.  
Some Pandas functions need the index to be set to a specific column (like `df.plot()`) or produce a specifically set index (like `df.groupby()`), but in the end it is just an additional column and there is no fundamental need for an index in a dataframe-structure.

---

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 1:11pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/11 "2021-12-08T13:11:29Z")

</div>

That’s all true, but I guess one reason I was asking this question is to assess how people feel about this behaviour: if lots of people are coming to Julia from Python and are used to Pandas then they might switch to a different package for data-wrangling, which would be valuable information for me

---

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [December 8, 2021, 1:18pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/12 "2021-12-08T13:18:17Z")

</div>

Coming from Pandas, DataFrames.jl works great for me. I like the minimalistic approach and you get essentially all relevant Pandas features with DataFrames.jl plus additional packages like CSV.jl, ShiftedArrays, etc.  
One big strengh of DataFrames.jl is that you can use any custom array element type with good performance, not just the built-in standard ones like in Pandas.

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 8, 2021, 1:27pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/13 "2021-12-08T13:27:56Z")

</div>

I spent a month in Pandas build an application for wrangling. “wow, this is cool”

Then I hit memory stress (not Panda’s fault), so I thought “hmm, perhaps Julia would be better”

So I rebuilt it in Julia, my first time stressing DataFrames in any way “wow, this is so much better than Pandas”

But really, I don’t think it matters what Data Format you produce. Unless you are outputting 100Gb of data, consumers can just transform it.

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [December 8, 2021, 1:33pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/14 "2021-12-08T13:33:57Z")

</div>

> [@Eagertom](#):
>
> by multi-dimensional I mean something like netCDF data, where a 3D grid might be required to either store 3D data or to store 2D data for a range of variables.

Maybe I don’t understand the usecase. But wouldn’t a regular 3D `Array` work? You can also fill arrays with arbitrary elements, like tuples, or even custom structs, with full performance. What do Dataframes add to the mix?

---

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 1:36pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/15 "2021-12-08T13:36:14Z")

</div>

We would like it to be as user friendly as possible - that by making a query using one of several functions (depending on the type of request) the returned data type is well presented and ready to use without writing a personal converter. Otherwise, frankly, the user might as well write the whole connector - all it does is parse your arguments as a URL and call HTTP.

That’s reminded me of another thing I felt was a limitation of DataFrames (or possibly of Julia?): that I couldn’t split out the metadata and associate it with the DataFrame (like an object attribute in Python). Of course, I could create a struct which contained the DataFrame and the metadata, but that’s another aspect that you have to explain to the user. Returning everything in a well established type which achieves all of this would make our documentation less meaty (although I appreciate that’s basically passing the buck to someone else’s careful documentation)

---

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 1:40pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/16 "2021-12-08T13:40:38Z")

</div>

Generally, labelling. It’s not clear from a 3D array of data alone that a given axis is latitude and another is longitude. IIRC Julia arrays don’t allow type mixing, so a String couldn’t be used to describe the remaining values in a row/column; even if you forced this by overloading Array, I don’t think you can then easily select based on column names. I know from experience of using numpy to manipulate netCDF data before Pandas and Xarray were known to me that this way of life is a nightmare (primarily from a debugging standpoint)

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [December 8, 2021, 1:46pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/17 "2021-12-08T13:46:56Z")

</div>

@Eagertom what is the topology of your data set? Is it like a raster file or like a point set or like a collection of polygons? Depending on this information we can provide more directions of data structures available in the language.

---

<div class="post-metadata">

**Author:** ![Eagertom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eagertom/32/24157_2.png) [@Eagertom](https://discourse.julialang.org/u/Eagertom)\
**Post date:** [December 8, 2021, 2:03pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/18 "2021-12-08T14:03:24Z")

</div>

I’m not quite sure I understand the question but I think raster. The data we provide is all gridded by lat/lon/altitude i.e. spherical polar.

I can provide more info if you can clarify. Also our [GitHub page](https://github.com/meteomatics/python-connector-api/tree/ee96e565cc8c64d6c08c0f7f81be285427c45597) and [API documentation](https://www.meteomatics.com/en/api/response) provide examples of features I’d ultimately like to implement. I already have an implementation which produces DataFrames, but want to ensure that this is really the best way of serving up the data for Julia users.

Thanks for the links to your own repositories provided above by the way - they look well developed and promise to shed some light on this topic for me!

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [December 8, 2021, 2:56pm UTC](https://discourse.julialang.org/t/most-popular-tabular-multidimensional-data-types-in-julia/72766/19 "2021-12-08T14:56:03Z")

</div>

> [@Eagertom](#):
>
> Generally, labelling. It’s not clear from a 3D array of data alone that a given axis is latitude and another is longitude. IIRC Julia arrays don’t allow type mixing, so a String couldn’t be used to describe the remaining values in a row/column; even if you forced this by overloading Array, I don’t think you can then easily select based on column names.

It sounds like you want `AxisArrays` I think. It allows you to name dimensions and have primary keys. But people more familiar with spatial / geological data may have better advice.
