# \[ANN\] Arrow.jl 0.3 Release

**URL:** <https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727>\
**Category:** Data\
**Tags:** arrow\
**Created:** [October 4, 2020, 4:59am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727 "2020-10-04T04:59:25Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 4, 2020, 4:59am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/1 "2020-10-04T04:59:25Z")

</div>

A new 0.3 release has just been made for the [Arrow.jl](https://github.com/JuliaData/Arrow.jl) package.

This is a significant rewrite of the entire package from scratch, and it now lives under the JuliaData organization. With this release, Arrow.jl now fully implements the 1.0 version of the [apache arrow](https://arrow.apache.org/docs/format/Columnar.html) format in native Julia. More detailed support now includes:

- All primitive data types
- All nested data types
- Dictionary encodings and messages
- Extension types
- Streaming, file, record batch, and replacement and isdelta dictionary messages

It currently doesn’t include support for:

- Tensors or sparse tensors
- Flight RPC
- C data interface

Third-party data formats:

- csv and parquet support via the existing CSV.jl and Parquet.jl packages
- Other Tables.jl-compatible packages automatically supported (DataFrames.jl, JSONTables.jl, JuliaDB.jl, SQLite.jl, MySQL.jl, JDBC.jl, ODBC.jl, XLSX.jl, etc.)
- No current Julia packages support ORC or Avro data formats

This 0.3 release is meant as a “beta” release of the new rewritten code and we invite all to give it a try and [report any issues](https://github.com/JuliaData/Arrow.jl/issues/new) you may run into. Also feel free to post questions/issues in the #data slack channel.

The plan is to let the 0.3 help shake out any glaring issues in the rewritten code before doing an official 1.0 release. In the mean time, I’ll also be working on integrating the julia implementation into the official apache arrow repository.

For the really adventurous among you, I recorded a [90-minute video](https://youtu.be/EXMRIBjxWFo) doing a deep-dive into the Arrow.jl Julia implementation of the arrow format; it dives deep into the code and also gives some high-level ideas/uses for arrow data in general.

Cheers!

-Jacob

---

<div class="post-metadata">

**Author:** ![yurivish](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yurivish/32/307_2.png) [@yurivish](https://discourse.julialang.org/u/yurivish)\
**Post date:** [October 4, 2020, 7:28am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/2 "2020-10-04T07:28:09Z")

</div>

This looks great! I have a question - is it possible to dict-encode a single column only (e.g. if I have a categorical column of strings in an otherwise numeric table)?

The docs state that this is controlled by a single option for all columns but I’m curious if this a choice made in the package or by the standard itself.

Thanks again for working on this; Arrow support is a really useful addition to the Julia data ecosystem.

Edit: Looking at [the code](https://github.com/JuliaData/Arrow.jl/blob/bd364c12048977a84f36baedcb25817144b090ec/src/write.jl#L250) it seems that the underlying format supports choosing to dict-encode or not dict-encode individual columns.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 4, 2020, 5:42pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/3 "2020-10-04T17:42:18Z")

</div>

Can we get at least one small benchmarking plot?

pretty please

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 4, 2020, 6:28pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/4 "2020-10-04T18:28:19Z")

</div>

Of what exactly? Comparison reading a file vs. csv or something? Comparison vs. other language implementations?

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [October 4, 2020, 6:33pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/5 "2020-10-04T18:33:17Z")

</div>

Yea I was interested between other file formats.

But perhaps since you say this is a pure Julia implementation of the library, might also be good to compare it to other languages as well. That’s less for people on this board though, more for people who find Arrow and then hopefully happen to see that the Julia implementation is one of the fastest!

---

<div class="post-metadata">

**Author:** ![cstjean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cstjean/32/1444_2.png) [@cstjean](https://discourse.julialang.org/u/cstjean)\
**Post date:** [October 5, 2020, 1:23am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/6 "2020-10-05T01:23:10Z")

</div>

Cool! Is there any future to Feather.jl, or should we just switch to Arrow? It says

> **NOTE:** Feather V1 has been deprecated by Apache Arrow in favor of Feather V2, which is just the Arrow IPC format written to disk. A complete rewrite of Arrow.jl is actively being worked on which will support reading and writing Feather V2… Currently Feather V2 will not be recognized as a valid feather file by this package.

But it hasn’t been worked on since several months it seems.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 5, 2020, 5:02am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/7 "2020-10-05T05:02:22Z")

</div>

Sorry for the slow response; yes, you can definitely dict encode single columns. The Arrow.jl package provides a wrapper type `DictEncode` which you can use to wrap your existing array and it will be dict encoded, or if your array is already a PooledArray/CategoricalArray type, it will be dict-encoded automatically (specifically, any array type that implements `DataAPI.refarray` and `DataAPI.refpool` apis). But yeah, simplest to dict-encode a single column is to just wrap it like `df.col1 = Arrow.DictEncode(df.col1)`

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 5, 2020, 5:04am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/8 "2020-10-05T05:04:30Z")

</div>

It’s a good question. I think Feather.jl is still technically alive, but I’m not exactly sure why. As far as I can tell, it’s just a very limited subset of the arrow format, on-disk only, and only really supported in C++, R, python, and Julia; and for the non-Julia implementations, they just use C++ arrow under the hood anyway.

My guess is it will be retired at some point and people will be encouraged to switch to arrow.

---

<div class="post-metadata">

**Author:** ![yurivish](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yurivish/32/307_2.png) [@yurivish](https://discourse.julialang.org/u/yurivish)\
**Post date:** [October 5, 2020, 5:26am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/9 "2020-10-05T05:26:36Z")

</div>

Thanks for the response! Cool, it’s great that this is already possible.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 5, 2020, 10:39pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/10 "2020-10-05T22:39:54Z")

</div>

Here’s a very quick comparison of a file with 70K rows of ints, floats, strings, and dates:

```julia
julia> @time f = CSV.File(file);
  0.004444 seconds (141.48 k allocations: 8.359 MiB)

julia> @time f2 = Arrow.Table(file2);
  0.000247 seconds (419 allocations: 26.016 KiB)

```

---

<div class="post-metadata">

**Author:** ![anon92994695](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@anon92994695](https://discourse.julialang.org/u/anon92994695)\
**Post date:** [October 5, 2020, 11:24pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/11 "2020-10-05T23:24:07Z")

</div>

Thank you! This is going to end up being one of my new favorite packages :D.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [October 6, 2020, 12:11am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/12 "2020-10-06T00:11:44Z")

</div>

Can you provide more details of your benchmark, please?  
Is “file2” a CSV file too? or a file with Arrow format?  
Could you also compare timings with data.table’s fread()?

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 6, 2020, 2:19am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/13 "2020-10-06T02:19:02Z")

</div>

No, `file2` is the same data as in `randoms.csv`, but in arrow format, so you’re seeing a comparison of CSV.jl reading a csv file, and Arrow.jl reading an arrow file. I don’t really have the time or desire to do detailed benchmarks right now vs. other languages/data formats, but if someone else would like to, feel free!

---

<div class="post-metadata">

**Author:** ![jablauvelt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jablauvelt/32/8792_2.png) [@jablauvelt](https://discourse.julialang.org/u/jablauvelt)\
**Post date:** [October 8, 2020, 11:44am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/14 "2020-10-08T11:44:31Z")

</div>

This is awesome and another big step forward for the Julia package ecosystem! Well done.

---

<div class="post-metadata">

**Author:** ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Post date:** [October 9, 2020, 7:43pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/15 "2020-10-09T19:43:35Z")

</div>

Hi,

Many thanks for your work on this - I am trying to incorporate Julia into an (“toy”) ETL of CSV → Arrow format since my goal is to the consume the data in a memory mapped form (via Julia) after the fact.

To this end I have imported the Arrow.jl package and read through the GitHub and essentially tried to recreate the above (using Julia v1.5) with the following code:

```julia
using Arrow , Tables , CSV , DataFrames
parfile = "C:\\Desktop\\juliaData\\csv\\gtscopyv4.csv"
arrowfile = "C:\\Desktop\\juliaData\\csv\\gtscopyv4.arrow"
#df = DataFrame!(CSV.File(parfile, header = false))
#io = open(parfile)
Arrow.write(arrowfile,CSV.File(parfile, header = false))

```

However, I consistently receive the following error:

```julia
ERROR: LoadError: `write` is not supported on non-isbits arrays
Stacktrace:
 [1] error(::String) at .\error.jl:33
 [2] write(::IOStream, ::CSV.File{false}) at .\io.jl:634
 [3] (::Base.var"#292#293"{CSV.File{false},Tuple{}})(::IOStream) at .\io.jl:396
 [4] open(::Base.var"#292#293"{CSV.File{false},Tuple{}}, ::String, ::Vararg{String,N} where N; kwargs::Base.Iterators.Pairs{Union{},Union{},Tuple{},NamedTuple{(),Tuple{}}}) at .\io.jl:325
 [5] open(::Function, ::String, ::String) at .\io.jl:323
 [6] write(::String, ::CSV.File{false}) at .\io.jl:396
 [7] top-level scope at c:\Users\doliver\Documents\JuliaRepo\TestCode\ArrowTest.jl:10
 [8] include(::String) at .\client.jl:457

```

The file itself contains some strings, ints:

![image](https://global.discourse-cdn.com/julialang/original/3X/e/a/ea9afcbbeb313380be662c2584e3895740f7f678.png)

Am I doing anything obviously wrong? I do note that this has been referenced in some earlier issue posts that appear to be resolved.

Apologies if this is the wrong place to be posting - this happens to be my first post here.

Regards

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [October 9, 2020, 7:48pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/16 "2020-10-09T19:48:29Z")

</div>

Hey @djholiver, sorry you’re having trouble. From the looks of it, your Arrow.jl version isn’t quite up to date. You can check this by doing `] st` which will list all the packages in your environment w/ their versions. You’ll want to make sure your Arrow.jl is 0.3; anything earlier won’t work the same.

---

<div class="post-metadata">

**Author:** ![anon69491625](https://avatars.discourse-cdn.com/v4/letter/a/aeb1de/32.png) [@anon69491625](https://discourse.julialang.org/u/anon69491625)\
**Post date:** [October 10, 2020, 4:43pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/17 "2020-10-10T16:43:08Z")

</div>

do you have any examples of using a python script to access an IN MEMORY Arrow.jl structure? We were going to use python to read a csv into memory on an inhouse server ( 256 gb ram) and use feather to offer up the data as an in memory structure. Happy to use Arrow.jl instead.

---

<div class="post-metadata">

**Author:** ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Post date:** [October 12, 2020, 1:35pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/18 "2020-10-12T13:35:39Z")

</div>

Hi @quinnj, thank you for your rapid response - I had switched off the laptop for the weekend (a rarity) so didn’t catch it in time.

I had ran Pkg.update(“Arrow”) before working with this initially, however, it didn’t actually update until I did a full Pkg.update() this morning.

Now that it has updated, I am able to do your example. This opens up myriad opportunities for query and analytical needs: my next steps involve reducing the data into separate query and index spaces - then binding this to a Julia service for on - demand data extraction.

What an amazing piece of code; kudos to you.

Regards

---

<div class="post-metadata">

**Author:** ![djholiver](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/djholiver/32/50470_2.png) [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Post date:** [November 19, 2020, 10:58pm UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/19 "2020-11-19T22:58:33Z")

</div>

Hi,

I’ve been making a lot of progress with this package and now have working pocs of a data query layer that reads directly from an arrow file, producing a parameter-based filtered output via an exported function.

What I want to do next is pass the filtered rows and a subset of the columns to another compute process(es). I am intending to define the column subset before calling the filter result, but I’m unsure on how best to proceed: should I wrap both the query and compute modules in a web service and do the communication over http, or would Arrow.write allow me to do IPC with zero copy etc.?

How would I write the latter? Would I call a while loop over the rows produced by the query layer from the compute layer(s)? My end goal is to have a distributed, load balanced query and compute approach.

Regards,

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [November 20, 2020, 4:49am UTC](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727/20 "2020-11-20T04:49:47Z")

</div>

Yeah, I think passing the data via IPC/on disk would work well. Once you perform the query on arrow data, you would just write the result via `Arrow.write("datadir/result.arrow", query_result)`.

The compute process could utilize [`FileWatching.watch_folder`](https://docs.julialang.org/en/v1.6-dev/stdlib/FileWatching/#FileWatching.watch_folder) to watch the data directory `datadir` for new query results to be written out. Once a new query is written out, it would just call `Arrow.Table("datadir/result.arrow"` to get the query result arrow data.

[Next page](https://discourse.julialang.org/t/ann-arrow-jl-0-3-release/47727.md?page=2)
