# Using JuliaDB to create larger than memory datasets and work with them?

**URL:** <https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724>\
**Category:** General Usage\
**Created:** [December 16, 2018, 8:36pm UTC](https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724 "2018-12-16T20:36:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [December 16, 2018, 8:36pm UTC](https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724/1 "2018-12-16T20:36:31Z")

</div>

JuliaDB docs explain how to perform some basic operations out-of-core, with data larger than memory.  
[http://juliadb.org/latest/manual/out-of-core.html](http://juliadb.org/latest/manual/out-of-core.html)  
But it seems it can only load large data and produce small results.

How can I get large results too (not only the input)?  
I’m using Julia 1.0.2 on Windows 10.

Imagine I want to do something like this:

```julia
using DataFrames
N=3
myDT = DataFrame(group = repeat('A':'C',outer=N), x = 1:(3*N) ) # create a dataframe
myDT.y = myDT.x .* rand(3*N) # add a new column z
myDT[myDT.group .== 'A', :y] = 0 # Replace y values when group == 'A'

```

but with much larger N, too large to fit on memory. (Or for example create two large matrices, multiply them and save the result).

How can I do it with JuliaDB for N larger than 10^9 and save it on disk?

I’ve tried

```julia
using JuliaDB
N=10^9
table((group = repeat('A':'C',outer=N), x = 1:(3*N) ))

```

but it consumes all my RAM and produces the error  
`ERROR: OutOfMemoryError()`

---

<div class="post-metadata">

**Author:** ![Noel\_Araujo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noel_araujo/32/4862_2.png) [@Noel\_Araujo](https://discourse.julialang.org/u/Noel_Araujo)\
**Post date:** [October 14, 2019, 11:00pm UTC](https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724/2 "2019-10-14T23:00:05Z")

</div>

Folks, I now that is late, but, any news on this post ?? 🙃

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [October 14, 2019, 11:55pm UTC](https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724/3 "2019-10-14T23:55:49Z")

</div>

AFAICT, this is not at all a JuliaDB issue. Try just running `repeat('A':'C',outer=10^9)` on a system that has only 4GB of RAM; it will likely throw the exact same error, because it allocates ~10GB of memory. That data must first be materialized before it even gets to a JuliaDB function.

Maybe the way to handle this is to make an iterator and have `table` support “unrolling” iterables while it writes them to disk; I’m not sure if that’s currently supported, but I doubt it’d be a hard PR.

---

<div class="post-metadata">

**Author:** ![Noel\_Araujo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/noel_araujo/32/4862_2.png) [@Noel\_Araujo](https://discourse.julialang.org/u/Noel_Araujo)\
**Post date:** [October 15, 2019, 12:55am UTC](https://discourse.julialang.org/t/using-juliadb-to-create-larger-than-memory-datasets-and-work-with-them/18724/4 "2019-10-15T00:55:45Z")

</div>

> Maybe the way to handle this is to make an iterator and have `table` support “unrolling” iterables while it writes them to disk; I’m not sure if that’s currently supported, but I doubt it’d be a hard PR.

Exactly this point of iterates that I was wandering if someone already did a PR, or if someone have a hack to share.
