# Multicollinearity and GLM

**URL:** <https://discourse.julialang.org/t/multicollinearity-and-glm/71340>\
**Category:** Statistics\
**Tags:** glm\
**Created:** [November 11, 2021, 5:25pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340 "2021-11-11T17:25:25Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![rikh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rikh/32/204104_2.png) [@rikh](https://discourse.julialang.org/u/rikh)\
**Post date:** [November 11, 2021, 5:25pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/1 "2021-11-11T17:25:25Z")

</div>

Below, I show some generated data for a linear regression. With these features (also known as variables, covariates or predictors) `A, B, C, D` and `E`, I aim to predict an outcome `Y`. Would the data shown below be considered multicollinear in the sense that it could become problematic for linear regressions?

 ![image](https://global.discourse-cdn.com/julialang/original/3X/2/7/27271d090c3b5abebaf6e7c8357c2f27d5a940b9.png)

I would say yes, and I read that Bayesian models can handle collinear data well, so I expected a huge difference between a Bayesian and Frequentist model. However, I compared a Bayesian to a Frequentist model and they gave the same outcomes, see the figures below. Therefore, I concluded that the Frequentist model did not have any issues with the collinearity.

Might this be because `lm` from GLM uses QR decomposition? Or, is my data not correlated enough? When would `GLM.lm` start showing huge variances as mentioned in [Wasserman’s lecture notes](https://www.stat.cmu.edu/~larry/=stat401/lecture-17.pdf)?

 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/d/5de7328f29d4542c1b26be03e348f2d96018bc73.png)

 ![image](https://global.discourse-cdn.com/julialang/original/3X/b/d/bde8fde64a2cda4d9c6e4c4d5f728201d5b8932c.png)

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 11, 2021, 5:38pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/2 "2021-11-11T17:38:44Z")

</div>

> [@rikh](#):
>
> Or, is my data not correlated enough?

Your data is not correlated enough. Multi-collinearity will only become a problem with estiamtion if its is very very high, i.e. within the margin of error for QR decomposition. A correlation of `.82` is not that.

---

<div class="post-metadata">

**Author:** ![rikh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rikh/32/204104_2.png) [@rikh](https://discourse.julialang.org/u/rikh)\
**Post date:** [November 11, 2021, 5:46pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/3 "2021-11-11T17:46:43Z")

</div>

> [@pdeffebach](#):
>
> Your data is not correlated enough.

You mean between the variables? Dormann et al. ([2012](https://doi.org/10.1111/j.1600-0587.2012.07348.x)) talk about degraded performance from correlation coefficients between variables of |r| \> 0.7. Maybe `Statistics.cor` is very different from r. I’ll look into that now.

EDIT: Nope that’s not it. Pearson’s r it is and `Statistics.cor` calculates the Pearson correlation too.

EDIT2: The correlations between the variables are definitely above 0.7:

```julia
julia> cor(df.D, df.E)
0.7750338235759653

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [November 11, 2021, 5:58pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/4 "2021-11-11T17:58:49Z")

</div>

I’m not familiar with the issues studied in the paper linked, but I’ve never heard of anyone in econometrics discuss |r| \> .7 being a problem.

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [November 11, 2021, 7:42pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/5 "2021-11-11T19:42:46Z")

</div>

[https://github.com/ericqu/LinearRegression.jl](https://github.com/ericqu/LinearRegression.jl) has a test for collinearity built in. Collinearity in linear regression models means that the coefficients will be estimated imprecisely. If the priors counter the particular imprecision, then Bayesian methods will help. But, if the priors don’t add information in the dimensions where it’s lacking, they won’t help much.

---

<div class="post-metadata">

**Author:** ![rikh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rikh/32/204104_2.png) [@rikh](https://discourse.julialang.org/u/rikh)\
**Post date:** [November 11, 2021, 9:29pm UTC](https://discourse.julialang.org/t/multicollinearity-and-glm/71340/6 "2021-11-11T21:29:44Z")

</div>

> [@mcreel](#):
>
> If the priors counter the particular imprecision, then Bayesian methods will help. But, if the priors don’t add information in the dimensions where it’s lacking, they won’t help much.

That makes sense. Thanks a lot!
