# Convert a (numerically) non-pd matrix into pd matrix

**URL:** <https://discourse.julialang.org/t/convert-a-numerically-non-pd-matrix-into-pd-matrix/24059>\
**Category:** General Usage\
**Created:** [May 10, 2019, 4:44am UTC](https://discourse.julialang.org/t/convert-a-numerically-non-pd-matrix-into-pd-matrix/24059 "2019-05-10T04:44:40Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mr.Robot](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mr.robot/32/8052_2.png) [@Mr.Robot](https://discourse.julialang.org/u/Mr.Robot)\
**Post date:** [May 10, 2019, 4:44am UTC](https://discourse.julialang.org/t/convert-a-numerically-non-pd-matrix-into-pd-matrix/24059/1 "2019-05-10T04:44:40Z")

</div>

## Problem

I am trying to compute a linear regression model that predicts overall delay of flights with respect to distance, carrier and some other predictors. I decided to use [SweepOperator.jl](https://github.com/joshday/SweepOperator.jl) because large number of observations (more than 40 million).

However, when I formed the Gram matrix, it is not positive definite (the snapshot of its eigenvalues are shown below), which is the case because of the numerical reasons. Applying sweep operator to this matrix will result in a lot of `NaN` in the estimated \hat{\beta}.

I do not know what I can do to make it positive definite to that it is pd (with some guarantee).

 ![image](https://global.discourse-cdn.com/julialang/original/3X/5/0/506366e0021662fe6752c97e3aa532cacd852adf.png)

## What I Have Done

I choose first 1 million observations and form a Gram matrix, which is also non-pd.

I (somehow) add an identity matrix to it and (somehow) make it pd. Since the entries of Gram matrix is extremely large, this seems to solve the problem. In fact, when I compared the results given by `sweep!()` and `sklearn.linear_model.LinearRegression()`, they are essentially the same.

However, I am not sure if this is a generally acceptable way and whether there is any rationale behind this.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [May 10, 2019, 7:57am UTC](https://discourse.julialang.org/t/convert-a-numerically-non-pd-matrix-into-pd-matrix/24059/2 "2019-05-10T07:57:27Z")

</div>

AFAIK there are specific algorithms for online updating of OLS that are fairly robust, eg

```nohighlight
@article{gentleman1974algorithm,
  title={Algorithm as 75: Basic procedures for large, sparse or weighted linear least problems},
  author={Gentleman, W Morven},
  journal={Journal of the Royal Statistical Society. Series C (Applied Statistics)},
  volume=23,
  number=3,
  pages={448--454},
  year=1974,
  publisher={JSTOR}
}

```

---

<div class="post-metadata">

**Author:** ![pistacliffcho](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pistacliffcho/32/8518_2.png) [@pistacliffcho](https://discourse.julialang.org/u/pistacliffcho)\
**Post date:** [May 10, 2019, 8:49am UTC](https://discourse.julialang.org/t/convert-a-numerically-non-pd-matrix-into-pd-matrix/24059/3 "2019-05-10T08:49:37Z")

</div>

In fact, what you have done is essentially recreated _ridge regression_, although it’s usually presented as lambda \* identity matrix, with the parameter lambda usually chosen by something like cross-validation. This creates a downward bias in the estimated coefficients _but_ often provides better out of sample estimates, especially for high dimensional problems.
