# Any Julia's equivalent to R's packages mcgv or mixed-effects models larger than memory?

**URL:** <https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865>\
**Category:** Statistics\
**Created:** [June 21, 2018, 3:52pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865 "2018-06-21T15:52:18Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [June 21, 2018, 3:52pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/1 "2018-06-21T15:52:19Z")

</div>

Hello.

I’m interested in fitting mixed effects models with large dataset that don’t fit on memory.  
R’s lme4 it’s too slow and doesn’t work if the dataset is large (a fraction of your RAM).  
speedglm and mgcv are a little bit faster but still have problems.

I’ve decided to move to Julia to try to find a better option.  
Mixedmodels.jl is like lme4.

Does Julia have something more like mgcv, faster and able to run Generalized additive mixed models?  
Or something able to fit mixed-effects models with datasets of around 50GB (on a computer with 12GB)?  
I mean not loading everything on memory, automatically streaming to disk as necessary.

Another option would be Spark or Flink, they work with very large datasets but I think they don’t have any implementation of mixed-effects models.

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [June 22, 2018, 5:50pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/2 "2018-06-22T17:50:56Z")

</div>

You may want to check [https://github.com/linkedin/photon-ml](https://github.com/linkedin/photon-ml)

I haven’t used it myself but they claim to be able to work with very large data sets.

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [June 22, 2018, 5:54pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/3 "2018-06-22T17:54:07Z")

</div>

By the way, although `MixedModels.jl` is similar in design to `lme4` (not surprising given my involvement in both projects) it is much more careful of the storage usage and generally much faster.

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 18, 2018, 6:14pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/4 "2018-11-18T18:14:02Z")

</div>

Unfortunatelly it doesn’t support random effect models, just fixed effects.

What about OnlineStats.jl or JuliaDB, is it possible to use them to fit random effect models?

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [November 18, 2018, 9:50pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/5 "2018-11-18T21:50:56Z")

</div>

In sas, this can be done but pretty slowly. In r, sparklyr can do regressions. For fixed effects, you can add dummy group variables by yourself. Not sure about out of memory random effect model fitting.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [November 18, 2018, 10:40pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/6 "2018-11-18T22:40:37Z")

</div>

Try looking into OnlineStats.jl. It can fit GLMs

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 19, 2018, 12:25am UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/7 "2018-11-19T00:25:28Z")

</div>

Any thoughts on making some of the `absorb` features from `reghdfe` available in `MixedModels.jl`?

---

<div class="post-metadata">

**Author:** ![Juan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juan/32/7657_2.png) [@Juan](https://discourse.julialang.org/u/Juan)\
**Post date:** [November 19, 2018, 12:55am UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/8 "2018-11-19T00:55:21Z")

</div>

I couldn’t find any example on how to run regressions with random effects (repeated measures) with OnlineStats.

I have opened a thread with an example and some benchmarks

> [@GLM is slow on large datasets. Using OnlineStats for regressions? MixedModels?](https://discourse.julialang.org/t/glm-is-slow-on-large-datasets-using-onlinestats-for-regressions-mixedmodels/17695):
>
> Hello. I’m planning to move from R to Julia and doing some tests about how to properly deal with large datasets and do simple tasks like regressions or survival analysis. I’ve done a benchmark with R (microbenchmark) for the regressions. N ← 3000 x1 ← rep(1:N, N) x2 ← rep(1:N, each = N) x3 ← sqrt(rep(1:N^2)) x1x2 ← x1x2 gg ← rep(1:5, each=N^2/5) y ← 1-2x1+3x2+0.5x1x2+rnorm(N^2)+x3\*rnorm(N^2) dat ← data.frame(y,x1,x2,x1x2,x3, gg) dat2 ← cbind(1,x1,x2,x1x2,x3,gg) lm(y ~ x1 + x2 + x1…

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [November 19, 2018, 4:16pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/9 "2018-11-19T16:16:26Z")

</div>

I don’t know what that is. I’ve never used Stata

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 19, 2018, 4:41pm UTC](https://discourse.julialang.org/t/any-julias-equivalent-to-rs-packages-mcgv-or-mixed-effects-models-larger-than-memory/11865/10 "2018-11-19T16:41:28Z")

</div>

As far as I can tell its an algorithm described in [Correia (2017): A Feasible Estimator for Linear Models with Multi-Way Fixed Effects](http://scorreia.com/research/hdfe.pdf) and outlined [here](http://scorreia.com/software/reghdfe/).

I recently had to implement a daily time series across 4 years of daily data and 110 regional units. Memory was a constraint even with `lme4`, though I didn’t get a chance to try `MixedModels`. The regression was instant with `reghdfe`, so I’m curious what the prospects are for incorporation into MixedModels, or even better, GLM.
