# \[ANN\] LinRegOutliers: a Julia package for detecting outliers in linear regression

**URL:** <https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280>\
**Category:** Package Announcements\
**Tags:** statistics, regression\
**Created:** [September 25, 2020, 6:10pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280 "2020-09-25T18:10:38Z")\
**Posts on this page:** 5\
**Page:** 2

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [August 13, 2021, 1:40pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/21 "2021-08-13T13:40:41Z")

</div>

Thanks for this package, which looks very useful. I’m interested in exploring possible outliers in a sample of a single random variable, rather than in a regression context. Would this be possible with the package? I have tried setting a formula where the dependent variable is regressed on a constant, and this seems to work in some cases, but not others (see example below). So, is this a reasonable idea for checking for outliers of a single random variable, when using the package? If so, of the methods the package provides, are there recommendations for this usage case? Thanks!

```julia
julia> using LinRegOutliers

julia> reg = createRegressionSetting(@formula(y ~ 1), hbk);

julia> hs93(reg)
Dict{Any, Any} with 3 entries:
  "outliers" => [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
  "t" => -3.53755
  "d" => [17.4311, 18.145, 18.502, 17.0741, 17.9665, 17.9665, 19.3944, 18.502, 17…

julia> smr98(reg)
ERROR: ArgumentError: Distance matrix should be symmetric.

```

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [August 16, 2021, 8:31am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/22 "2021-08-16T08:31:01Z")

</div>

When you set a formula using `y ~ 1`, a regression model of `y = constant + epsilon` is estimated and it is still a regression model. I think it is more convenient to use single variable tools in that situation. `smr98` is based on a cluster analysis on standardized tuples (yhat, residuals), I think the problem is all of the `yhat` values are the same in your example (or something else that I can’t cover at first sight). Some of the other algorithms will work in the univariate case but I wouldn’t say this is a recommended method.

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [August 16, 2021, 10:45am UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/23 "2021-08-16T10:45:54Z")

</div>

Thanks. I realized that I needed to read more about the methods, rather than just use them without checking suitability of a method for this usage case.

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [December 29, 2022, 5:15pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/24 "2022-12-29T17:15:01Z")

</div>

After 2 years since the first post, we got much attention to the package. As a re-announcement we are happy to revise our package to v0.8.16, with the latest implementation of the Quantile Regression estimator. Here is the latest list of the algorithms & estimators:

- Ordinary Least Squares, Weighted Least Squares, Basic diagnostics
- Hadi & Simonoff (1993)
- Kianifard & Swallow (1989)
- Sebert & Montgomery & Rollier (1998)
- Least Median of Squares
- Least Trimmed Squares
- Minimum Volume Ellipsoid (MVE)
- MVE & LTS Plot
- Billor & Chatterjee & Hadi (2006)
- Pena & Yohai (1995)
- Satman (2013)
- Satman (2015)
- Setan & Halim & Mohd (2000)
- Least Absolute Deviations (LAD)
- Quantile Regression Parameter Estimation (quantileregression)
- Least Trimmed Absolute Deviations (LTA)
- Hadi (1992)
- Marchette & Solka (2003) Data Images
- Satman’s GA based LTS estimation (2012)
- Fischler & Bolles (1981) RANSAC Algorithm
- Minimum Covariance Determinant Estimator
- Imon (2005) Algorithm
- Barratt & Angeris & Boyd (2020) CCF algorithm
- Atkinson (1994) Forward Search Algorithm
- BACON Algorithm (Billor & Hadi & Velleman (2000))
- Hadi (1994) Algorithm
- Chatterjee & Mächler (1997)
- Summary

Thank you the Julia community for the attention!

---

<div class="post-metadata">

**Author:** ![jbytecode](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jbytecode/32/17719_2.png) [@jbytecode](https://discourse.julialang.org/u/jbytecode)\
**Post date:** [June 18, 2026, 7:50pm UTC](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280/25 "2026-06-18T19:50:20Z")

</div>

I have updated the package to [v0.11.8](https://github.com/jbytecode/LinRegOutliers/commit/98d6791ab9c6e22bbab1db4e06cc83270e213675 "pump v0.11.8"). In the new patch release, memory allocations have been reduced. The `lts()` function (Least trimmed squared regression) has shorter times than the R counterpart which is based on native code (`ltsreg()` function in MASS package).

[Previous page](https://discourse.julialang.org/t/ann-linregoutliers-a-julia-package-for-detecting-outliers-in-linear-regression/47280.md?page=1)
