# \[ANN\] LinRegOutliers: a Julia package for detecting outliers in linear regression

**URL:** <https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280>\
**Category:** Package Announcements\
**Tags:** statistics, regression\
**Created:** [September 25, 2020, 6:10pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280 "2020-09-25T18:10:38Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 25, 2020, 6:10pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/1 "2020-09-25T18:10:38Z")

</div>

Hi there,

This is an open invitation for the community. LinRegOutliers is a Julia package for detecting outliers in linear regression. In its current state, the package contains some pioneering algorithms that have taken place in the literature. The scope of the package is direct and indirect (robust) outlier detection methods in linear regression and multivariate location and scale estimators.

Implementations of the some missing methods assigned to our newly joined friends. There are more!

If you want to have contributions, you are welcome.

Please visit the repo at [LinRegOutliers GitHub repository](https://github.com/jbytecode/LinRegOutliers)

---

<div class="post-metadata">

**Author:** ![Paul\_Soderlind](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paul_soderlind/32/1753_2.png) [@Paul\_Soderlind](https://discourse.julialang.org/u/Paul_Soderlind)\
**Post date:** [September 25, 2020, 8:31pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/2 "2020-09-25T20:31:44Z")

</div>

This looks very promising. Just one question: the examples all use `@formula(y ~ x1 + x2 + x3)` etc. Is this required, or would `(y,x)` work? (where `y` is a vector and `x` a matrix)

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [September 25, 2020, 8:37pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/3 "2020-09-25T20:37:43Z")

</div>

I made a PR to GLM.jl with Cook’s Distance that I never felt the need to finish (shame on me), perhaps this is a better place for it to live if you don’t have it already.

[https://github.com/JuliaStats/GLM.jl/pull/368](https://github.com/JuliaStats/GLM.jl/pull/368)

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [September 25, 2020, 8:40pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/4 "2020-09-25T20:40:01Z")

</div>

Not that this isn’t a good feature, but you know that you can generate formulas programmatically, right? if you have a data frame `df`, you could actually do

```julia
term("y") ~ sum(term.(names(df, Not("y")))

```

to get many covariates at the same time.

---

<div class="post-metadata">

**Author:** ![Paul\_Soderlind](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paul_soderlind/32/1753_2.png) [@Paul\_Soderlind](https://discourse.julialang.org/u/Paul_Soderlind)\
**Post date:** [September 25, 2020, 8:52pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/5 "2020-09-25T20:52:35Z")

</div>

Agreed, this works. It would still be nice if there were methods for `(y,x)`. It’s as simple as it gets and often as good as the more complicated alternatives.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 26, 2020, 6:49am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/6 "2020-09-26T06:49:29Z")

</div>

We have cooks() in src/diagnostics.jl as a tool in an outlier detection algorithm.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 26, 2020, 6:53am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/7 "2020-09-26T06:53:10Z")

</div>

You are right, since the models are linear, the response vector and a design matrix are required and a tuple of (y, X) is enough. However, construction of design matrices takes time as they are not always simple, because of dummy variables, handling variable slopes and intercepts, interactions etc. The FormulaTerm or @formula works well with lm() in GLM. After all, this is the standard way of describing a regression model for both Julia GLM and R users.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 26, 2020, 11:28am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/8 "2020-09-26T11:28:30Z")

</div>

> [@jbytecode](#):
>
> After all, this is the standard way of describing a regression model for both Julia GLM and R users.

This is true, but occasionally an `y, X` interface can be useful too.

If it already exists, perhaps making it part of the API would be nice.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 27, 2020, 6:00pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/9 "2020-09-27T18:00:29Z")

</div>

Hi @Tamas_Papp

Thank to multiple dispatch feature of Julia, we may have many methods for an algorithm.

@Paul_Soderlind opened a new issue for this, I offered to have second definitions of all methods that accepts data in the form of (y, X) tuple. He said he could help whenever he had time.

Thank you 🙂

---

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [September 27, 2020, 7:18pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/10 "2020-09-27T19:18:50Z")

</div>

BTW, most of what you import from GLM is actually re-exports of stuff from StatsModels.jl (everything except `lm`). And for version of StatsModels 0.6 or greater, the `ModelFrame`/`ModelMatrix` interface is discouraged in favor of doing something like

```julia
concrete_f = apply_schema(f, schema(f, data))
y,X = modelcols(concrete_f, data)

```

I don’t think it’s a huge deal to use that interface (and it might simplify some of the code actually), by storing `concrete_f` in the regression setting struct. You can get just the response or design matrix by calling `modelcols(f.lhs, data)` or `modelcols(f.rhs, data)` (respectively).

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 27, 2020, 7:28pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/11 "2020-09-27T19:28:46Z")

</div>

Okay, I note this, thank you. If you want to join, you are welcome.

---

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [September 27, 2020, 7:37pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/12 "2020-09-27T19:37:01Z")

</div>

Yeah if I have the time I’ll put in a PR 🙂 In the mean time, the docs for StatsModels are (I hope) comprehensive enough to get started… (here’s the [deep dive](https://juliastats.org/StatsModels.jl/stable/internals/#The-lifecycle-of-a-@formula-1))

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 27, 2020, 7:49pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/13 "2020-09-27T19:49:33Z")

</div>

Okay, to be honest, I’m too focused on writing new methods in the package. I care the literature on this subject to be production ready and to be complete in Julia. My friends and people like you would help on the performance improvements, code optimizations, aesthetically improving, etc.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [September 28, 2020, 6:46pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/14 "2020-09-28T18:46:13Z")

</div>

We have now an issue about `(y, X)` in [issue(14)](https://github.com/jbytecode/LinRegOutliers/issues/14). Of course, the needs of the community must be taken into account.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [October 7, 2020, 3:12pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/15 "2020-10-07T15:12:33Z")

</div>

Hi folks!

`(X, y)` style multiple dispatch is implemented except a single method and the next revision will include

```julia
algorithm(X::Matrix, y::Vector)

```

beyond the original version

```julia
algorithm(setting::RegressionSetting)

```

for all. Fyi.

Best!

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [January 5, 2021, 6:32pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/16 "2021-01-05T18:32:07Z")

</div>

Hey guys,

I am happy to announce that the research paper of our Julia package LinRegOutliers has just been published in Journal of Open Source Software.

Here is the citation info and link:

[https://joss.theoj.org/papers/10.21105/joss.02892](https://joss.theoj.org/papers/10.21105/joss.02892)

Satman et al., (2021). LinRegOutliers: A Julia package for detecting outliers in linear regression. Journal of Open Source Software, 6(57), 2892, [Journal of Open Source Software: LinRegOutliers: A Julia package for detecting outliers in linear regression](https://doi.org/10.21105/joss.02892)

Do not forget to cite us! 🥰

---

<div class="post-metadata">

**Author:** ![Dario\_Sarra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dario_sarra/32/9338_2.png) [@Dario\_Sarra](https://discourse.julialang.org/u/Dario_Sarra)\
**Post date:** [March 16, 2021, 3:53pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/17 "2021-03-16T15:53:02Z")

</div>

I generally use linear mixed models from the MIxedModels.jl package. Can these methods be used for those cases as well?

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [March 17, 2021, 5:25am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/18 "2021-03-17T05:25:01Z")

</div>

This package requires the model at hand is in form of y=X\beta + \epsilon and the estimation is based on least squares or maximum likelihood estimation. Implemented methods take the model as a @formula object or X and y where X is the design matrix and y is the response vector. As I know and correct me if I am wrong, the main estimation process in mixed models is generally handled by EM algorithm. I have never researched the literature on this topic, but there may be other outlier detection algorithms for such models.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [March 17, 2021, 8:33am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/19 "2021-03-17T08:33:19Z")

</div>

> [@jbytecode](#):
>
> there may be other outlier detection algorithms for such models

“Outlier” is not a precisely defined term, and deterministic binary classification of whether something is an “outlier” is not a particularly robust method for inference. The advantage is, of course, that it is fast.

It is usually better to assume a mixture distribution (eg adding something with a fat tail), or start with something like a Student-t in the first place for the errors. This, of course, is better suited to Bayesian inference, or EM.

---

<div class="post-metadata">

**Author:** ![Dario\_Sarra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dario_sarra/32/9338_2.png) [@Dario\_Sarra](https://discourse.julialang.org/u/Dario_Sarra)\
**Post date:** [March 17, 2021, 12:09pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/20 "2021-03-17T12:09:12Z")

</div>

I have to clarify that I am definitely not an expert in statistics, so I might be missing some basic understanding of the issue. Anyway, MixedModels package use maximum likelihood to fit the formula. Does this mean I can use your approach to detect outliers then?

[Next page](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280.md?page=2)
