# Reading large-columned data using Feather.jl is too slow

**URL:** <https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377>\
**Category:** Data\
**Tags:** question, package\
**Created:** [May 28, 2020, 8:56pm UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377 "2020-05-28T20:56:18Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean/32/15263_2.png) [@Jean](https://discourse.julialang.org/u/Jean)\
**Post date:** [May 28, 2020, 8:56pm UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/1 "2020-05-28T20:56:18Z")

</div>

Hi, I am working with large-columned data sets. One small example is the dimension of (1452, 66584) whose data size is about 2GB. When it converted to a feather format, its size was down to 778MB. The problem is Feather.jl is unexpectedly slow for first reading, taking fairly large memory allocation, so I failed several times in ACF for memory issue and unexpectedly fast after that. Here is the output I succeeded in the following local machine:

```julia
Julia> using Feather

Julia> @time al=Feather.read("DO_gm_ofa_unadj_alpr_ch1.feather");
1850.519921 seconds (17.67 G allocations: 363.156 GiB, 1.80% gc time)

julia> @time al=Feather.read("DO_gm_ofa_unadj_alpr_ch1.feather");
  3.102820 seconds (17.18 M allocations: 575.066 MiB, 7.20% gc time)

```

```julia
Julia> versioninfo()
Julia Version 1.0.5
Commit 3af96bcefc (2019-09-09 19:06 UTC)

Platform Info:
OS: Linux (x86_64-pc-linux-gnu)
CPU: Intel(R) Xeon(R) CPU E5-2630 v3 @ 2.40GHz
WORD_SIZE: 64
LIBM: libopenlibm
LLVM: libLLVM-6.0.0 (ORCJIT, haswell)

```

This file is not the only to work with; it is one of the files I jointly work with. Do you have any idea to read large-columned data fast?

Thanks.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 29, 2020, 1:23am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/2 "2020-05-29T01:23:54Z")

</div>

Do you have lots of missings and lots of `String` columns?

Well you could try [JDF.jl](https://github.com/xiaodaigh/JDF.jl) if interop with Python and R is not a big priority. It is generally quite fast for me (I am biased cos I developed it).

Parquet.jl’s reader and writer are not the best in terms of performance atm.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [May 29, 2020, 1:39am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/3 "2020-05-29T01:39:18Z")

</div>

This will be at least partially true for any file format reader. I see this with CSV, SASLib, etc…

Although I will say that 1850 seconds is a bit of a stretch! I read files in that are much larger than 2GB (again, CSV or SASLib) and I never see those kind of times. 2-3 minutes tops.

---

<div class="post-metadata">

**Author:** ![Jean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean/32/15263_2.png) [@Jean](https://discourse.julialang.org/u/Jean)\
**Post date:** [May 29, 2020, 2:25am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/4 "2020-05-29T02:25:06Z")

</div>

The datasets are float64 and no missing. When compared with large-rowed data, my case is fairly slow.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [May 29, 2020, 2:28am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/5 "2020-05-29T02:28:23Z")

</div>

In that case JDF.jl will perform very well and is suited to your use case.

---

<div class="post-metadata">

**Author:** ![MarkovChains](https://avatars.discourse-cdn.com/v4/letter/m/f04885/32.png) [@MarkovChains](https://discourse.julialang.org/u/MarkovChains)\
**Post date:** [June 27, 2020, 10:48pm UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/6 "2020-06-27T22:48:32Z")

</div>

Let us know if JDF.jl solves the issue!

---

<div class="post-metadata">

**Author:** ![Jean](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jean/32/15263_2.png) [@Jean](https://discourse.julialang.org/u/Jean)\
**Post date:** [June 28, 2020, 3:24am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/7 "2020-06-28T03:24:29Z")

</div>

I tried to use JDF.jl but it wasn’t satisfiable. In my group, my colleague developed a new pkg ‘Helium. jl’ to fix this issue. It will soon be released.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [June 28, 2020, 4:07am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/8 "2020-06-28T04:07:08Z")

</div>

> [@Jean](#):
>
> tried to use JDF.jl but it wasn’t satisfiable

You mean you couldn’t install it or the feature are not up to scratch, in terms of speed or usage? I am interested to know what are the failing if you can be so kind to volunteer your time to answer my question.

I try to make JDF better.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [June 28, 2020, 4:09am UTC](https://discourse.julialang.org/t/reading-large-columned-data-using-feather-jl-is-too-slow/40377/9 "2020-06-28T04:09:07Z")

</div>

> [@Jean](#):
>
> Helium. jl

I can’t find the repo at all. Is it on github?
