# Disk based data manipulation framework needed

**URL:** <https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303>\
**Category:** Data\
**Tags:** data\
**Created:** [October 7, 2017, 6:13am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303 "2017-10-07T06:13:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 7, 2017, 6:13am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/1 "2017-10-07T06:13:21Z")

</div>

SAS is big in the corporate world. I work in the finance industry and SAS is still quite big. To be honest, using SAS is a pain. SAS can’t even syntax highlight its own language properly. But it’s still there because it has one trick – disk-based data manipulation and associated algorithms.

Ten years ago when I introduced R to my workplace, people were skeptical – you can’t load a large dataset in R and manipulate it like in SAS. That’s because in R the dataset needs to be loaded into memory and at that time the largest laptop only had 4G of RAM. Today, 32G RAM laptops are becoming the norm but still I can’t load really large (50G) datasets into RAM.

I think Julia and R can replace SAS by implementing disk based data manipulation as a first class citizen. Also can most algorithms works once the data becomes disk based? If not then Julia still can’t replace SAS, because most of SAS’s algorithms (e.g. proc glm) works off disk based data

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 7, 2017, 6:26am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/2 "2017-10-07T06:26:27Z")

</div>

I am not sure I understand what you are saying. If you are advocating the development of libraries that have functionality like SAS, the best way to do that is to start working on one.

If you are asking about disk-based data access: it is very easy to do for large data using `mmap`. I am in the process of working on a project that involves this, and will do a blog post soon, but the principle is very simple: map a file and and array, and from then on just access your data with `[]`.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 7, 2017, 6:33am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/3 "2017-10-07T06:33:09Z")

</div>

Nice. I have started working on some functions that work on feather files stored manually as chunks. Hoepfully it will turn into a package later on. I will look into mmap seems prettt cool.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 7, 2017, 6:37am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/4 "2017-10-07T06:37:52Z")

</div>

Looking forward to your blog post. I represent the proportion of people who have never heard of mmap before.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [October 7, 2017, 6:43am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/5 "2017-10-07T06:43:13Z")

</div>

Have you looked at JuliaDB.jl?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 7, 2017, 6:50am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/6 "2017-10-07T06:50:10Z")

</div>

Thanks!! I thought JuliaDB was for connection to databases. Didn’t realise it had persistent data storage capability. Looks very close to what I need. Will do the research.

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [October 7, 2017, 7:02am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/7 "2017-10-07T07:02:00Z")

</div>

Also, there will soon be [OnlineStats](https://github.com/joshday/OnlineStats.jl/) integration in JuliaDB ([https://github.com/JuliaComputing/JuliaDB.jl/pull/75](https://github.com/JuliaComputing/JuliaDB.jl/pull/75)) which would help building algorithms on top of it. Take a look at [SparseRegression](https://github.com/joshday/SparseRegression.jl), for an example.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [October 7, 2017, 7:13am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/8 "2017-10-07T07:13:42Z")

</div>

It’s a bit tricky, as there’s a “JuliaDB” organisation for connecting to databases, and then there’s the unrelated “JuliaDB.jl” package…

---

<div class="post-metadata">

**Author:** ![tim.holy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tim.holy/32/52_2.png) [@tim.holy](https://discourse.julialang.org/u/tim.holy)\
**Post date:** [October 7, 2017, 12:06pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/9 "2017-10-07T12:06:57Z")

</div>

@xiaodai, Julia has some amazing tools for big data. One example is the ability to do lazy transformations of large arrays. For example, let’s imagine you have a 10TB 4d array stored as an NRRD file, and you want to take the square root of each element and swap dimensions 3 and 4. This could easy take a couple of hours using other tools, and would involve writing out another disk file in the process. In Julia it only takes a few microseconds and can be done “in memory”:

```julia
using FileIO, MappedArrays
A = load("bigfile.nrrd")
C = PermutedDimsArray(mappedarray(sqrt, A), (1,2,4,3))

```

That’s because all the operations here are lazy (“virtual”) and are computed on-demand. You can pass these lazy arrays to visualization code, etc, and as long as it’s all been written against our generic AbstractArray interface it should all Just Work.

Of course Julia also supports eager computation (which would be `permutedims(sqrt.(A), (1,2,4,3))`), but for big data lazy is very nice.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 7, 2017, 12:16pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/10 "2017-10-07T12:16:48Z")

</div>

I hope to be able to learn more about these and be able to introduce this to the masses. It’s not something that I’ve seen and the syntax looks a bit different to the type programming I am used to e.g. R data.frame, data.table.

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [October 7, 2017, 12:23pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/11 "2017-10-07T12:23:37Z")

</div>

It’s also worth mentioning packages wrapping SQL engines, like SQLite. I know SAS users often rely on `proc sql` because it’s faster than the standard `data` step, so that should make sense to them. Of course that requires writing SQL instructions.

I think @davidanthoff has also been working on a SQL backend to Query.jl, which would essentially allow you to run the same query against a data frame or against a SQL database depending on your needs.

---

<div class="post-metadata">

**Author:** ![ChrisRackauckas](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/chrisrackauckas/32/77_2.png) [@ChrisRackauckas](https://discourse.julialang.org/u/ChrisRackauckas)\
**Post date:** [October 8, 2017, 3:20am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/12 "2017-10-08T03:20:33Z")

</div>

I don’t think it’s been mentioned in the thread, but the term you’re looking for is out-of-core.

> **[External memory algorithm](https://en.wikipedia.org/wiki/Out-of-core_algorithm)**
>
> In computing, external memory algorithms or out-of-core algorithms are algorithms that are designed to process data that are too large to fit into a computer's main memory at once. Such algorithms must be optimized to efficiently fetch and access data stored in slow bulk memory (auxiliary memory) such as hard drives or tape drives, or when memory is on a computer network. External memory algorithms are analyzed in the external memory model.
> External memory algorithms are analyzed in an idealized...

JuliaDB does out-of-core through Dagger.jl, and databases like SQL do this as well like @nalimilan says.

But one of the important things with Julia is distinguishing between the representation of data and the API. Using generic functions with dispatch, the same API can apply to many different “backends” which handle the data differently. So you may want to look at interfaces like this (Query.jl, DataStreams.jl, IterableTables.jl, etc) to mix the choices depending on the circumstance, but using the same code.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 31, 2017, 10:37am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/13 "2017-10-31T10:37:59Z")

</div>

I did a short writeup here:  
[https://tpapp.github.io/post/large-ragged-dataset-julia/](https://tpapp.github.io/post/large-ragged-dataset-julia/)  
Does not go into much detail, but the libraries I made public are much better documented. Hope you find this useful.

FWIW, once data is ingested into a binary format and mmapped, I find that I can process a 100 GB dataset in a few minutes with a reasonably recent computer (even a laptop) with an SSD. The key is almost-linear access, random access is of course much worse.

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [October 31, 2017, 11:01am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/14 "2017-10-31T11:01:44Z")

</div>

There isn’t a minimal subset of the dataset anywhere for trying out your code?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 31, 2017, 11:25am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/15 "2017-10-31T11:25:10Z")

</div>

I will create one soon if that would help.

---

<div class="post-metadata">

**Author:** ![sdanisch](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sdanisch/32/1406_2.png) [@sdanisch](https://discourse.julialang.org/u/sdanisch)\
**Post date:** [October 31, 2017, 11:35am UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/16 "2017-10-31T11:35:37Z")

</div>

I just wrote a very similar blog post: [Drawing 2.7 billion Points in 10s | by Simon Danisch | HackerNoon.com | Medium](https://medium.com/@sdanisch/drawing-2-7-billion-points-in-10s-ecc8c85ca8fa)  
🙂 Not sure how on topic this is, but it’s at least disc based!

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 31, 2017, 1:51pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/17 "2017-10-31T13:51:02Z")

</div>

Interesting writeup, thanks! Regarding `write`: AFAICT there is no simple `write(::IO, ::T)` where `isbits(T)` even in `master`, so I submitted a PR:  
[https://github.com/JuliaLang/julia/pull/24234](https://github.com/JuliaLang/julia/pull/24234)  
but since you know much more about the internals, maybe you could suggest an improvement or make another PR that does this.

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [October 31, 2017, 1:59pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/18 "2017-10-31T13:59:54Z")

</div>

If anyone is interested, [Feather.jl](https://github.com/JuliaData/Feather.jl) is already quite useful for working with memory mapped data via [this PR](https://github.com/JuliaData/Feather.jl/pull/54). I already use it that way quite routinely (also feather is a really wonderful format). I really should talk to @quinnj about getting that merged, but I’ve been happily using my fork and have mostly forgotten about it.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [October 31, 2017, 2:03pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/19 "2017-10-31T14:03:45Z")

</div>

I did look at `Feather.jl`, and found two problems with it:

1. I need to know the data size in advance (which requires another pass),
2. AFAICT types are restricted to what Feather supports (is this correct?)

---

<div class="post-metadata">

**Author:** ![ExpandingMan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/expandingman/32/866_2.png) [@ExpandingMan](https://discourse.julialang.org/u/ExpandingMan)\
**Post date:** [October 31, 2017, 2:08pm UTC](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303/20 "2017-10-31T14:08:36Z")

</div>

> [@Tamas\_Papp](#):
>
> AFAICT types are restricted to what Feather supports (is this correct?)

Yes, that is certainly true. If you have need of custom datatypes, Feather is definitely not for you. In those cases I use JLD, but I rarely have much need to store large amounts of data of custom types.

> [@Tamas\_Papp](#):
>
> I need to know the data size in advance (which requires another pass),

For writing you mean? Yes, that seems to be a limitation as well. At least in my case I usually “write once, read millions of times”.

[Next page](https://discourse.julialang.org/t/disk-based-data-manipulation-framework-needed/6303.md?page=2)
