# Lazily fetch/load data into a DataFrame

**URL:** <https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283>\
**Category:** General Usage\
**Created:** [November 15, 2023, 9:48pm UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283 "2023-11-15T21:48:34Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 15, 2023, 9:48pm UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283/1 "2023-11-15T21:48:34Z")

</div>

What’s the best way to lazily fetch/load data into a DataFrame, only if/when needed? For example, if a user calls a function `use_data_set1()`, then I want to fetch/load `data_set1` from some remote URL and make it available for the rest of that user’s session. However, if the user never calls a function that relies on `data_set1`, I don’t want to fetch & load it. If the data is fetched/loaded once, I don’t want to fetch/load it again on subsequent function calls that utilize that data…

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 15, 2023, 10:33pm UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283/2 "2023-11-15T22:33:11Z")

</div>

There are Julia packages which formalize this process more, that other people can mention. But here is a quick and dirty solution using a global `Dict` to store data sets.

```julia
julia> using DataFrames

julia> global DATA_DICT = Dict();

julia> function load_state_data()
           url = "url_to_state_data"
           DataFrame(state = 1:50, value = 100 .* rand(10))
       end;

julia> function load_county_data()
           url = "url_to_county_data"
           DataFrame(county = 1:10, value = rand(10))
       end;

julia> function analyze_county()
           county_data = get!(load_county_data, DATA_DICT, :county_data)
           println(county_data)
       end
analyze_county (generic function with 1 method)

julia> function analyze_state()
           county_data = get!(load_county_data, DATA_DICT, :county_data)
           println(county_data)
       end
analyze_state (generic function with 1 method)

julia> analyze_county()
10×2 DataFrame
 Row │ county value    
     │ Int64 Float64  
─────┼──────────────────
   1 │ 1 0.759557
   2 │ 2 0.383699
   3 │ 3 0.851332
   4 │ 4 0.928275
   5 │ 5 0.433502
   6 │ 6 0.691074
   7 │ 7 0.619731
   8 │ 8 0.475289
   9 │ 9 0.347691
  10 │ 10 0.163557

julia> DATA_DICT
Dict{Any, Any} with 1 entry:
  :county_data => 10×2 DataFrame…

```

---

<div class="post-metadata">

**Author:** ![rdavis120](https://avatars.discourse-cdn.com/v4/letter/r/b5a626/32.png) [@rdavis120](https://discourse.julialang.org/u/rdavis120)\
**Post date:** [November 16, 2023, 12:50am UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283/3 "2023-11-16T00:50:33Z")

</div>

It’s not necessarily a Julia solution but I use [duckdb](https://duckdb.org/docs/archive/0.9.2/api/julia) to create a table view of multiple files using glob file names:  
CREATE VIEW users AS SELECT \* FROM ‘/\*/test.parquet’;

Alternatively for urls you can use the [httpfs](https://duckdb.org/docs/archive/0.9.2/extensions/httpfs) extension for json or parquet files. For example:  
SELECT \* FROM read\_parquet(‘s3://bucket/\*.parquet’);

You could then get the output of the query as a DataFrame if you need to do further processing in memory.

---

<div class="post-metadata">

**Author:** ![mthelm85](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mthelm85/32/224164_2.png) [@mthelm85](https://discourse.julialang.org/u/mthelm85)\
**Post date:** [November 17, 2023, 5:43pm UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283/4 "2023-11-17T17:43:38Z")

</div>

Is something like the following any better (or is it worse?) than using the global `Dict`?

```julia
using DataFrames

const data = Ref{Union{Nothing, DataFrame}}(nothing)

function get_data()
    if isnothing(data[])
        data[] = DataFrame(a=rand(10),b=rand(10))
    end
end

```

This way I have other functions that call `get_data()`, but if it’s already been called, it won’t do anything.

---

<div class="post-metadata">

**Author:** ![rdavis120](https://avatars.discourse-cdn.com/v4/letter/r/b5a626/32.png) [@rdavis120](https://discourse.julialang.org/u/rdavis120)\
**Post date:** [November 17, 2023, 10:35pm UTC](https://discourse.julialang.org/t/lazily-fetch-load-data-into-a-dataframe/106283/5 "2023-11-17T22:35:50Z")

</div>

I’m not sure if you already considered this but you might want to look at either package for memoization like [Memoization.jl](https://www.juliapackages.com/p/memoization)
