# How to add/query a dict that contains tens of millions of items?

**URL:** <https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886>\
**Category:** General Usage\
**Tags:** question\
**Created:** [September 15, 2023, 10:03am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886 "2023-09-15T10:03:58Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![leonchen2012](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leonchen2012/32/51130_2.png) [@leonchen2012](https://discourse.julialang.org/u/leonchen2012)\
**Post date:** [September 15, 2023, 10:03am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886/1 "2023-09-15T10:03:58Z")

</div>

I’m using a dict with struct key and struct value. Size of each key value pair is 0.5~3k bytes. The problem is that it can be accumulated to tens of millions of key-value pairs, which would cause OOM on my computer.  
I use it as a cache, i.e. just add and query. I know disk based key-value store can do it. But the LevelDB and RocksDB libs of Julia were quite outdated. They don’t even pass Julia compiling.  
Is there any other way I can try?

Thanks

---

<div class="post-metadata">

**Author:** ![barucden](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/barucden/32/26154_2.png) [@barucden](https://discourse.julialang.org/u/barucden)\
**Post date:** [September 15, 2023, 11:12am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886/2 "2023-09-15T11:12:42Z")

</div>

If it’s supposed to be a cache, then maybe you don’t need to keep all items the whole time? If that’s the case, [GitHub - JuliaCollections/LRUCache.jl: An implementation of an LRU Cache in Julia](https://github.com/JuliaCollections/LRUCache.jl) might help.

---

<div class="post-metadata">

**Author:** ![stephancb](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stephancb/32/14243_2.png) [@stephancb](https://discourse.julialang.org/u/stephancb)\
**Post date:** [September 15, 2023, 11:32am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886/3 "2023-09-15T11:32:26Z")

</div>

[Here](https://charlesleifer.com/blog/lsm-key-value-storage-in-sqlite3/) is a discussion of SQLite’s [lsm1](https://www.sqlite.org/cgi/src/dir?name=ext/lsm1) extension. It perhaps can handle your use case. To manage everything with Julia, the [SQLite.jl](https://juliadatabases.org/SQLite.jl/stable/) package is up-to-date, well maintained and quite efficient.

A similar approach using Rust (instead of Julia) is presented [here](https://blog.helsing.ai/lsmlite-rs-rust-bindings-for-sqlites-lsm1-storage-engine-30d710083062).

---

<div class="post-metadata">

**Author:** ![leonchen2012](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leonchen2012/32/51130_2.png) [@leonchen2012](https://discourse.julialang.org/u/leonchen2012)\
**Post date:** [September 18, 2023, 9:31am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886/4 "2023-09-18T09:31:37Z")

</div>

> [@stephancb](#):
>
> Here

Unfortunately, I have to keep all items for later usage. Those items will be saved to disk and loaded for next runnings.

---

<div class="post-metadata">

**Author:** ![leonchen2012](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/leonchen2012/32/51130_2.png) [@leonchen2012](https://discourse.julialang.org/u/leonchen2012)\
**Post date:** [September 18, 2023, 9:32am UTC](https://discourse.julialang.org/t/how-to-add-query-a-dict-that-contains-tens-of-millions-of-items/103886/5 "2023-09-18T09:32:29Z")

</div>

Thanks. I’ll try it.
