# Benchmarks of Various Formats for Tabular Data

**URL:** <https://discourse.julialang.org/t/benchmarks-of-various-formats-for-tabular-data/50489>\
**Category:** Data\
**Tags:** benchmark, input-output\
**Created:** [November 20, 2020, 11:13am UTC](https://discourse.julialang.org/t/benchmarks-of-various-formats-for-tabular-data/50489 "2020-11-20T11:13:17Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![lungben](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lungben/32/12314_2.png) [@lungben](https://discourse.julialang.org/u/lungben)\
**Post date:** [November 20, 2020, 11:13am UTC](https://discourse.julialang.org/t/benchmarks-of-various-formats-for-tabular-data/50489/1 "2020-11-20T11:13:18Z")

</div>

Hi,

here is a Pluto.jl notebook for benchmarking of the read and write performance, as well as file sizes of various formats for tabular data.

> <https://gist.github.com/lungben/7d967eb5058bbe5708bb08fb1aeb2815>

The following formats / packages are compared:

- CSV via [GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other delimited files in the Julia programming language](https://github.com/JuliaData/CSV.jl)
- JSON via [GitHub - JuliaData/JSONTables.jl: JSON3.jl + Tables.jl](https://github.com/JuliaData/JSONTables.jl)
- Zipped CSV via [GitHub - fhs/ZipFile.jl: Read/Write ZIP archives in Julia](https://github.com/fhs/ZipFile.jl)
- JDF via [GitHub - xiaodaigh/JDF.jl: Julia DataFrames serialization format](https://github.com/xiaodaigh/JDF.jl)
- Parquet via [GitHub - JuliaIO/Parquet.jl: Julia implementation of Parquet columnar file format reader](https://github.com/JuliaIO/Parquet.jl)
- Apache Arrow via [GitHub - apache/arrow-julia: Official Julia implementation of Apache Arrow](https://github.com/JuliaData/Arrow.jl)
- Excel (xlsx) via [GitHub - felipenoris/XLSX.jl: Excel file reader and writer for the Julia language.](https://github.com/felipenoris/XLSX.jl)
- SQLite via [GitHub - JuliaDatabases/SQLite.jl: A Julia interface to the SQLite library](https://github.com/JuliaDatabases/SQLite.jl)

I always used the default configuration of each package, i.e. multithreading or compression is only used if it is switched on by default.

On my machine, Arrow is fastest, followed by JDF (note that Arrow is not compressed per default by JDF is).

 ![timings_100000](https://global.discourse-cdn.com/julialang/original/3X/e/4/e4da23b193cbe8efe5a7f3ce628ef65b5c9fca8a.png)

 ![filesizes_100000](https://global.discourse-cdn.com/julialang/original/3X/2/8/282a17edd8d233f7e078f12b0c518ab07ddb15ed.png)

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [November 21, 2020, 4:14am UTC](https://discourse.julialang.org/t/benchmarks-of-various-formats-for-tabular-data/50489/2 "2020-11-21T04:14:38Z")

</div>

Very interesting! Thanks for sharing!
