# Really, really big arrays and OutOfMemoryError()

**URL:** <https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [November 21, 2019, 11:54pm UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366 "2019-11-21T23:54:08Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![clibassi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/clibassi/32/12856_2.png) [@clibassi](https://discourse.julialang.org/u/clibassi)\
**Post date:** [November 21, 2019, 11:54pm UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/1 "2019-11-21T23:54:08Z")

</div>

Hi all,

I’m a relative newbie to Julia and trying to do some record linkage in Julia. In that process, I need to bring two large datasets (1.8million rows and 800K rows respectively) into an array to loop through and do comparisons of the strings in those datasets. I began this process by attempting to initialize an empty array that can store all the comparisons (ie a 1.8m x 800k array). When I attempt to do this I get an OutOfMemoryError(). I have tried initializing it at a few different sizes, and only after reducing the size considerably am I able to create the array (examples below). I’m wondering if what I’m doing is just foolish or if there is another type of array I should be using to accomplish my goal. Many thanks!

 ![image](https://global.discourse-cdn.com/julialang/original/3X/2/5/25541dfc0935dacaffdcabe9aeeb42b129330820.png)

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [November 22, 2019, 12:00am UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/2 "2019-11-22T00:00:00Z")

</div>

How much memory do you have? `18000*900000*3=49GB`

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [November 22, 2019, 12:03am UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/3 "2019-11-22T00:03:33Z")

</div>

What comparison do you need to do? I think the solution has to be context dependent.

---

<div class="post-metadata">

**Author:** ![heliosdrm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/heliosdrm/32/3851_2.png) [@heliosdrm](https://discourse.julialang.org/u/heliosdrm)\
**Post date:** [November 22, 2019, 12:39am UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/4 "2019-11-22T00:39:24Z")

</div>

Perhaps you may consider _not_ creating such a large array. You might read the strings of your datasets from the source files one at each time (or in reasonably sized batches), make the comparison(s), and write the results in an output file incrementally, without keeping the whole data sets in memory.

File IO is pretty efficient in Julia, so don’t be wary of reading and writing files frequently. In cases like this it may be much faster than loading and handling big datasets in memory.

---

<div class="post-metadata">

**Author:** ![clibassi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/clibassi/32/12856_2.png) [@clibassi](https://discourse.julialang.org/u/clibassi)\
**Post date:** [November 24, 2019, 4:50pm UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/5 "2019-11-24T16:50:37Z")

</div>

I’m running on a machine with 24GB RAM. And the comparison I am trying to do is Levenshtein distance. I think I may just approach it by doing some blocking to cut down the number of comparisons by a lot.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [November 24, 2019, 4:57pm UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/6 "2019-11-24T16:57:47Z")

</div>

Old question of mine covering similar ground (but not well discoverable from the title):

[Why do these two functions benchmark the same?](https://discourse.julialang.org/t/why-do-these-two-functions-benchmark-the-same/12831)

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [November 24, 2019, 5:44pm UTC](https://discourse.julialang.org/t/really-really-big-arrays-and-outofmemoryerror/31366/7 "2019-11-24T17:44:34Z")

</div>

Your immediate problem of not having enough memory may be solved using [memory mapping](https://docs.julialang.org/en/v1/stdlib/Mmap/#Mmap.mmap), but the question is if you really need all those distances at the same time and whether or not you can use problem specific knowledge to avoid calculating them at all. Performance might be a problem though. For us to be able to give advice in that direction, we’re going to need some more information about what you’re doing with that data 🙂
