# What fileformat to use to load data for high performance computing

**URL:** <https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786>\
**Category:** Machine Learning\
**Created:** [July 11, 2017, 9:37am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786 "2017-07-11T09:37:32Z")\
**Posts on this page:** 18\
**Page:** 2

<div class="post-metadata">

**Author:** ![ScottPJones](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/scottpjones/32/146_2.png) [@ScottPJones](https://discourse.julialang.org/u/ScottPJones)\
**Post date:** [July 20, 2017, 8:54pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/21 "2017-07-20T20:54:01Z")

</div>

In a previous life, I was one of the principal system architects of an OLTP oriented database system (all proprietary, used heavily in the healthcare / insurance industries), which is also used to store the data from ESA’s Gaia project, so it’s interesting to me to see the rather different design choices made for something like CarbonData.

Thanks for the link!

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 20, 2017, 9:03pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/22 "2017-07-20T21:03:55Z")

</div>

At the end of the presentation they talk about future things and on one of those is to also have a row based format for fast data ingestion (like in Streaming) and then convert that to column based for OLAP/analystics so you kind of get a mixed OLTP/OLAP story.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [July 21, 2017, 3:39am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/23 "2017-07-21T03:39:35Z")

</div>

Thank you for a discussion. My feeling is that the state is not really nice, as there is no easy solution. I personally like feather files, as column structure makes a lot of sense to me.

I would like to ask Anthony, if he would object against support for reading Feather files from streaming sources. I have tried to “hack” the solution by reading the entire stream to memory first and then let it the existing parser parse the data. Although the change is tiny and conceptually does not make sense, it helps a lot when you are reading gzipped files (they nicely compress).

The connected question is, if someone knows how to produce feather files from Spark? I have not find any solution.

Thanks for the answers.

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 24, 2017, 11:15am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/24 "2017-07-24T11:15:02Z")

</div>

You might also want to check out [https://kudu.apache.org/](https://kudu.apache.org/)

It would be interesting to see a benchmark of Spark (SQLor dataFrames/dataTables API) on Parquet vs Spark on Carbondata vs Spark on Kudu. I’ve seen some benchmarks with some of these in it but not all and never Carbondata vs Kudu (I guess they are too new).

I think Kudu and Carbondata are quite close feature wise except that Kudu isn’t a public file format (on object storage cloud filesystem) but an object store.

---

<div class="post-metadata">

**Author:** ![Pevnak](https://avatars.discourse-cdn.com/v4/letter/p/2bfe46/32.png) [@Pevnak](https://discourse.julialang.org/u/Pevnak)\
**Post date:** [July 25, 2017, 9:43am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/25 "2017-07-25T09:43:14Z")

</div>

Can you conveniently load the data to julia?

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 25, 2017, 12:10pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/26 "2017-07-25T12:10:13Z")

</div>

Not yet. Spark.jl would need to expand the Spark API coverage and start supporting the Spark SQL/Dataframe/Dataset API of Spark 2.x instead of the low level RDD API of Spark1.x. Once that is in place you could read in all sorts of extra datasources supported by Spark including Carbondata & Kudu.  
Or for a Kudu specific low level RDD approach: Spark.jl would need to wrap KuduRDD.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [July 25, 2017, 12:58pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/27 "2017-07-25T12:58:26Z")

</div>

> [@Steven\_Sagaert](#):
>
> Spark.jl would need to expand the Spark API coverage and start supporting the Spark SQL/Dataframe/Dataset API of Spark 2.x instead of the low level RDD API of Spark1.x.

The basic support is already [there](https://github.com/dfdx/Spark.jl/blob/master/src/sql.jl), I just don’t have time to cover the rest. As far as I understand, supporting Kudu and Carbondata should be pretty easy, if you show me how to read these formats from Java (or Scala _without implicit conversions_, since they aren’t part of JNI and can’t be used from JavaCall), I’ll try to add corresponding functions to Spark.jl.

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 25, 2017, 1:13pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/28 "2017-07-25T13:13:38Z")

</div>

For Kudu see [Apache Kudu - Developing Applications With Apache Kudu](https://kudu.apache.org/docs/developing.html) and/or [Cloudera Blog](https://blog.cloudera.com/blog/2017/02/up-and-running-with-apache-spark-on-apache-kudu/)

for CarbonData: see the slides from this presentation [Home - Data + AI Summit 2022 | Databricks](https://spark-summit.org/east-2017/events/apache-carbondata-an-indexed-columnar-file-format-for-interactive-query-with-spark-sql/)

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [July 25, 2017, 2:55pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/29 "2017-07-25T14:55:05Z")

</div>

For Kudu I really need either Java code, or Scala without implicit conversions. For example this line:

```nohighlight
customersAppendDF.write.options(kuduOptions).mode("append").kudu

```

implicitly converts Spark’s `Dataset` into something Kudu-specific that has a `.kudu()` method. Since I’ve never worked with Kudu I don’t know what and when is converted, and don’t really have time to investigate it.

For Carbondata, generic `read_df` and `write_df` should work (just added them). Do you have an example of Carbondata file that I can test reading on?

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 25, 2017, 7:23pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/30 "2017-07-25T19:23:48Z")

</div>

You probably can do the insert/read/Write/update from KuduContext instead of via Dataframe

For the implicit method “kudu”  
I think I found it in the Git logs:

+++ b/java/kudu-spark/src/main/scala/org/apache/kudu/spark/kudu/package.scala  
@@ -32,7 +32,7 @@ package object kudu {  
\* Adds a method, `kudu`, to DataFrameWriter that allows writes to Kudu  
using  
\* the DataFileWriter  
\*/

- implicit class KuduDataFrameWriter(writer: DataFrameWriter) {
- 

```
 def kudu = writer.format("org.apache.kudu.spark.kudu").save

```

- }

- implicit class KuduDataFrameWriter(writer: DataFrameWriter) {
- def kudu = writer.format(“org.apache.kudu.spark.kudu”).save
- }  
}

so it’s a method on the package object kudu in java/kudu-spark/src/main/scala/org/apache/kudu/spark/kudu/package.scala.

I don’t have a carbondata file but you can read in data from another format like csv save it in carbondata and then read it back.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [July 25, 2017, 11:23pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/31 "2017-07-25T23:23:42Z")

</div>

> [@Steven\_Sagaert](#):
>
> def kudu = writer.format(“org.apache.kudu.spark.kudu”).save

This line is essential, thanks.

However, it seems like Kudu isn’t just a file format, but rather a distributed storage - all examples I’ve seen require specifying “kudu.master” and “kudu.table” options instead of file path. If this is correct, Kudu looks out of scope of this discussion (although may be in scope of Spark.jl). Have anybody used Kudu in practice to confirm or deny my assumption?

As for Carbondata, their integration with Spark breaks JSON support. I think I will wait until it gets more stable before including it into Spark.jl.

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 25, 2017, 11:30pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/32 "2017-07-25T23:30:59Z")

</div>

that’s correct : Kudu isn’t a file format but a distributed store which can serve as a datasource to Spark.

The reason I brought it up in this thread is because it seems to ly in between pure olap and pure oltp and Carbondata seems to also have some oltp features and I wanted to contrast it to that. Anyway from the point of Spark it’s pretty much all the same: a datasource and querying it is the same (SQL/Dataframe/Dataset).

---

<div class="post-metadata">

**Author:** ![Pevnak](https://avatars.discourse-cdn.com/v4/letter/p/2bfe46/32.png) [@Pevnak](https://discourse.julialang.org/u/Pevnak)\
**Post date:** [July 26, 2017, 4:44am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/33 "2017-07-26T04:44:00Z")

</div>

It seems like that the good file format is really missing for hpc is missing. Feather seems to be promising, but it breaks when you try to save and load unicode characters (we have issued bug report). Moreover, I have not find any feather library for Spark.  
Loading of Protobufs to Julia is surprisingly slow, because it seems like that the code is not type-stable.  
HDF5 (JLD) has poor support in Spark and it produces larger files than gziped feather, even when the compression is turned on.  
I am little bit desperate and do not see a good solution by now.

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 26, 2017, 7:56am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/34 "2017-07-26T07:56:37Z")

</div>

have you tried this [https://github.com/valiantljk/h5spark](https://github.com/valiantljk/h5spark)?

---

<div class="post-metadata">

**Author:** ![Pevnak](https://avatars.discourse-cdn.com/v4/letter/p/2bfe46/32.png) [@Pevnak](https://discourse.julialang.org/u/Pevnak)\
**Post date:** [July 26, 2017, 9:45am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/35 "2017-07-26T09:45:35Z")

</div>

No,  
but it only supports reading of HDF5, not writing. We need to go in the direction from preparing data in Spark and do the processing in Julia (TensorFlow).  
Tomas

---

<div class="post-metadata">

**Author:** ![Steven\_Sagaert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/steven_sagaert/32/29578_2.png) [@Steven\_Sagaert](https://discourse.julialang.org/u/Steven_Sagaert)\
**Post date:** [July 26, 2017, 10:52am UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/36 "2017-07-26T10:52:04Z")

</div>

That seems backwards to me but ok: just use Python in Spark and use [http://www.h5py.org/](http://www.h5py.org/) to write the data to HDF5.

---

<div class="post-metadata">

**Author:** ![dfdx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dfdx/32/120_2.png) [@dfdx](https://discourse.julialang.org/u/dfdx)\
**Post date:** [July 26, 2017, 1:29pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/37 "2017-07-26T13:29:32Z")

</div>

> [@Steven\_Sagaert](#):
>
> Anyway from the point of Spark it’s pretty much all the same: a datasource and querying it is the same (SQL/Dataframe/Dataset).

I haven’t tried it myself since I don’t have Kudu installed, but something like this should work:

1. Checkout latest master of Spark.jl.
2. In `~/.julia/v0.6/Spark/jvm/sparkjl/pom.xml`, uncomment dependencies related to Kudu.
3. Run `mvn clean package` from the same dir to rebuild main JAR.
4. Run from Julia:

```julia
spark = SparkSession()
df = read_json(spark, "path/to/sample.jsom")
options = Dict("mode" => "append", "kudu.master" => "kudu.master:7051", "kudu.table" => "test_table")
write_df(df; format="org.apache.kudu.spark.kudu", options=options)

```

Essentially, you can try out any Spark-compatible file format using this approach - just add required dependencies and specify correct format and options.

---

<div class="post-metadata">

**Author:** ![bchi](https://avatars.discourse-cdn.com/v4/letter/b/58956e/32.png) [@bchi](https://discourse.julialang.org/u/bchi)\
**Post date:** [December 1, 2018, 3:41pm UTC](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786/38 "2018-12-01T15:41:02Z")

</div>

Apache Arrow might be another option. Arrow and Parquet are merging code base.

[Previous page](https://discourse.julialang.org/t/what-fileformat-to-use-to-load-data-for-high-performance-computing/4786.md?page=1)
