# An example of Apache Arrow file?

**URL:** <https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299>\
**Category:** Data\
**Tags:** arrow\
**Created:** [March 31, 2021, 2:33pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299 "2021-03-31T14:33:31Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sami](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sami/32/20726_2.png) [@Sami](https://discourse.julialang.org/u/Sami)\
**Post date:** [March 31, 2021, 2:33pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/1 "2021-03-31T14:33:31Z")

</div>

Hello,

I would like to test Julia’s support for Apache Arrow with a big file that doesn’t fit to memory (RAM). I can’t find any example of such file, could you help me with this please?

I guess I need a file that is in Arrow IPC Format (Feather file format), version 2 (see for example [Feather File Format — Apache Arrow v9.0.0](https://arrow.apache.org/docs/python/feather.html)).

I would like to read the file into a dataframe like this: `df = DataFrame(Arrow.Table(the_big_file))`, as exemplified here: [User Manual · Arrow.jl](https://arrow.juliadata.org/dev/manual/)

Thanks for creating Julia 👍

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [March 31, 2021, 2:42pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/2 "2021-03-31T14:42:38Z")

</div>

I’d just find a large dataset and write it out. For example, the [nyc taxi dataset](https://www1.nyc.gov/site/tlc/about/tlc-trip-record-data.page) is a common example.

You’d download a couple files, then do something like:

```julia
Arrow.write("taxi.arrow",
    Tables.partitioner(csv_files) do file
        CSV.File(file)
    end
)

```

now you have a local arrow file called `"taxi.arrow"`. Hope that helps.

---

<div class="post-metadata">

**Author:** ![StatisticalMouse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/statisticalmouse/32/43370_2.png) [@StatisticalMouse](https://discourse.julialang.org/u/StatisticalMouse)\
**Post date:** [March 31, 2021, 2:58pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/3 "2021-03-31T14:58:44Z")

</div>

I believe the easiest location to obtain a lot of that would be [New York City Taxi and Limousine Commission (TLC) Trip Record Data - Registry of Open Data on AWS](https://registry.opendata.aws/nyc-tlc-trip-records-pds/)  
or possibly this [https://opendata.cityofnewyork.us/](https://opendata.cityofnewyork.us/)

A few months ago I tried a similar thing that didn’t work, but I tried parquet files as input, and I got those from an R package.

---

<div class="post-metadata">

**Author:** ![Sami](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sami/32/20726_2.png) [@Sami](https://discourse.julialang.org/u/Sami)\
**Post date:** [April 1, 2021, 2:35pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/4 "2021-04-01T14:35:02Z")

</div>

Thanks for the super fast replies @quinnj and @StatisticalMouse . I did something that combined your hints: I downloaded parquet files with R (by following this [Working with Arrow Datasets and dplyr • Arrow R Package](https://arrow.apache.org/docs/r/articles/dataset.html) ) and then tried to combine a big arrow file by `using Parquet; Arrow.write("taxi.arrow", Tables.partitioner(read_parquet(".")));`. These lines seem to crash Julia, I just get “Killed” message. I have Julia 1.6.0.

Do you have an idea why Julia crashes on this?

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [April 1, 2021, 3:00pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/5 "2021-04-01T15:00:41Z")

</div>

Hmmmm, not sure; it may be running out of memory. I don’t think Parquet.jl currently supports partitioned datasets, so I think it may be materializin the full parquet dataset in memory then trying to write it out to arrow memory.

---

<div class="post-metadata">

**Author:** ![StatisticalMouse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/statisticalmouse/32/43370_2.png) [@StatisticalMouse](https://discourse.julialang.org/u/StatisticalMouse)\
**Post date:** [April 1, 2021, 4:57pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/6 "2021-04-01T16:57:21Z")

</div>

That’s exactly what I tried earlier; it doesn’t work.

---

<div class="post-metadata">

**Author:** ![StatisticalMouse](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/statisticalmouse/32/43370_2.png) [@StatisticalMouse](https://discourse.julialang.org/u/StatisticalMouse)\
**Post date:** [April 1, 2021, 4:58pm UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/7 "2021-04-01T16:58:24Z")

</div>

I also have a recollection that there were two parquet reading packages.

---

<div class="post-metadata">

**Author:** ![Sami](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sami/32/20726_2.png) [@Sami](https://discourse.julialang.org/u/Sami)\
**Post date:** [April 22, 2021, 11:56am UTC](https://discourse.julialang.org/t/an-example-of-apache-arrow-file/58299/8 "2021-04-22T11:56:53Z")

</div>

I ended up doing the big arrow file with pyarrow, along with lines below:

```julia
with pa.output_stream("path/big.arrow") as sink:
    with pa.ipc.new_file(sink, schema) as writer:
        for arrowfile in glob.glob("path/to/files/*.arrow", recursive=False):
            with pa.input_stream(arrowfile) as source:
                with pa.ipc.open_file(source) as reader:
                    for i in range(0,reader.num_record_batches):
                        writer.write_batch(reader.get_batch(i))

```

That led to the another issue: [How well Apache Arrow’s zero copy methodology is supported?](https://discourse.julialang.org/t/how-well-apache-arrow-s-zero-copy-methodology-is-supported/59797)
