# Have confusion regarding Linear Regression

**URL:** <https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399>\
**Category:** New to Julia\
**Tags:** question, statistics, linear-regression\
**Created:** [January 28, 2022, 10:46pm UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399 "2022-01-28T22:46:03Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![mirror63](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@mirror63](https://discourse.julialang.org/u/mirror63)\
**Post date:** [January 28, 2022, 10:46pm UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/1 "2022-01-28T22:46:03Z")

</div>

Blood pressure is normally measured using cuff based methods. New contactless methods for blood pressure measurements rely on measuring time difference between two biomedical waveforms. This time difference is called pulse transit time or pulse arrival time (PAT). The relationship between systolic blood pressure and PAT is not clear but there are several potential options:  
BP = a+b_PAT + noise  
BP=c+d_ln(PAT) + noise

I am trying to do traditional linear gression, is it possible to create chains for the variables a,b,c,d and noise so I can show varianes of these data?

My confusion is that, as far as I know, for plots of the chain for a,b,c,d I can do it for Bayesian Linear Regression, using MCMC according to the source in [Linear Regression](https://turing.ml/dev/tutorials/05-linear-regression/) . However, for traditional linear regression I am confused about how to generate chain plots for the variable a,b,c,d. Is it possible to plot chain diagrams for simple linear regression?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [January 29, 2022, 6:56am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/2 "2022-01-29T06:56:59Z")

</div>

What is traditional linear regression? If you mean standard OLS estimated based on the closed form solution to the least squares problem, (X'X)^{-1}X'Y, then inference will normally be done based on the asymptotic distribution of the estimator, which again is available in closed form. There’s no sampling and therefore no chain.

This isn’t really a Julia question I suppose, you might want to consult a standard textbook that discussed this, such as Wooldridge’s Introductory Econometrics.

---

<div class="post-metadata">

**Author:** ![mirror63](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@mirror63](https://discourse.julialang.org/u/mirror63)\
**Post date:** [January 29, 2022, 7:15am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/3 "2022-01-29T07:15:12Z")

</div>

thank you so much. I am new to statistics (linear regression), and trying to solve problems in Julia as well. When I was trying to solve this, I had the same question cause it not make sense to me, cause i thought you can create chains for the lls/ols linear regression. Anyways, I guess you can create sample for Bayesian Linear Regression, and form chains using MCMC, which can later provide mean and variance for the posterior distribution. Thanks.

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [January 29, 2022, 7:54am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/4 "2022-01-29T07:54:28Z")

</div>

Traditional linear regression, as was noted above, leads to a point estimator. If the errors are normally distributed, then the small sample distribution of the estimator will also be normal. With non-normal errors, the asymptotic distribution will still be normal. However, the small sample distribution will not be so. Sometimes, bootstrapping is used in this context to explore the small sample distribution. In that context, it could make sense to represent the bootstrap samples using a chain. Here’s a code example that does that:

```julia
using Plots, Distributions, Statistics

# simple iid bootstrap
function bootstrap(data)
    n = size(data,1)
    resampled = similar(data)
    for i = 1:n
        j = rand(1:n)
        resampled[i,:] = data[j,:]
    end
    return resampled
end    

n = 50
reps = 1000
x = [ones(n) randn(n)]
β = [2.,-1.]
ϵ = rand(Chisq(3.),n) .- 1.5
y = x*β + ϵ
data = [y x]
bs = zeros(reps,2)
for i = 1:reps
    d = bootstrap(data)
    bs[i,:] = d[:,2:end] \ d[:,1]
end
plot(bs[:,2], labels=false)
q05 = quantile(bs[:,2], 0.05)
q95 = quantile(bs[:,2], 0.95)
hline!([q05], labels="q05")
hline!([q95], labels="q95")

```

---

<div class="post-metadata">

**Author:** ![mirror63](https://avatars.discourse-cdn.com/v4/letter/m/7ba0ec/32.png) [@mirror63](https://discourse.julialang.org/u/mirror63)\
**Post date:** [January 29, 2022, 8:22am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/5 "2022-01-29T08:22:12Z")

</div>

thank you so much, I have another confusion. In doing traditional linear regression, say OLS/LLS, we do not consider the error when finding the coefficients and y-intercept, can you explain why so?

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [January 29, 2022, 8:32am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/6 "2022-01-29T08:32:57Z")

</div>

The assumption is that the errors are not observed, so we need an estimator of the coefficients that depends only on the observed data.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [January 29, 2022, 11:51am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/7 "2022-01-29T11:51:18Z")

</div>

> [@mcreel](#):
>
> ```julia
> # simple iid bootstrap
> function bootstrap(data)
> n = size(data,1)
> resampled = similar(data)
> for i = 1:n
> j = rand(1:n)
> resampled[i,:] = data[j,:]
> end
> return resampled
> end 
> 
> ```

Isn’t this equivalent to the following one-liner:

```julia
bootstrap(data) = data[rand(1:end,end),:]

```

?

---

<div class="post-metadata">

**Author:** ![mcreel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcreel/32/30088_2.png) [@mcreel](https://discourse.julialang.org/u/mcreel)\
**Post date:** [January 29, 2022, 11:59am UTC](https://discourse.julialang.org/t/have-confusion-regarding-linear-regression/75399/8 "2022-01-29T11:59:37Z")

</div>

Yes, I noticed that, too, after I posted it. I wrote that function a long time ago… Probably, I should fix it to work with arrays of different sizes.
