# Why does GLM.jl fail to fit my logistic regression model but scikit learn and R has no issues? I think I found the reason

**URL:** <https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553>\
**Category:** Modelling & Simulations\
**Tags:** glm, r\
**Created:** [August 24, 2024, 9:45am UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553 "2024-08-24T09:45:23Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 24, 2024, 9:45am UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/1 "2024-08-24T09:45:23Z")

</div>

I have been trying to use GLM.jl to fit some logistic regression models and the fit failed. I don’t remember the exact error message but it was something about singularity.

That annoyed me for a couple of years but I never figured why because using RCall.jl was straight-forward enough and I just use it to fit the logistic regression.

Recently, I decided to look into the scikit-learn implementation and it occurred to me why that is the case!

The GLM.jl implementation using Julia’s powerful numerical abilities implemented the EM algorithms where as scikit just invoked a numerical solver that doesn’t use EM.

The EM must have some technical condition where the fit will fail while scikit-learn and R’s implementation will just treat it like any other numerical problem and try to find the best coefficients!

Admittedly the R’s coefficients for the problem I tried to solve contained some `NA` coefficients which I needed to discard but I actually prefer it to GLM.jl’s approach.

In Julia, is there an implementation of GLM that uses a solver like they do in R? Can’t seem to find one but I guess it’s easy to structure the problem as an optimisation problem and invoke a solver.

---

<div class="post-metadata">

**Author:** ![EOhneberg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/eohneberg/32/45704_2.png) [@EOhneberg](https://discourse.julialang.org/u/EOhneberg)\
**Post date:** [August 24, 2024, 11:22am UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/2 "2024-08-24T11:22:39Z")

</div>

Not an expert and don’t have an answer but I wonder whether it has anything to do with precision/scale. I had a similar problem before between R and Matlab. Turns out the default scale of Matlab was higher than in R resulting in singularity in R and a result in Matlab.

Edit: Agree that Julia should aim to obtain the same results as R when it comes to statistical analysis. Currently this doesn’t seem to be the case for other models also

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 24, 2024, 11:31am UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/3 "2024-08-24T11:31:11Z")

</div>

I highly doubt it. But I can try to experiment.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [August 24, 2024, 1:44pm UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/4 "2024-08-24T13:44:34Z")

</div>

> [@xiaodai](#):
>
> I have been trying to use [GLM.jl](https://juliahub.com/ui/Packages/General/GLM) to fit some logistic regression models and the fit failed. I don’t remember the exact error message but it was something about singularity.

If you can give a reproducible example and the exact error message, peopel can give you more helpful advice. Without that, it’s hard for this thread to go anywhere productive. [Please read: make it easier to help you](https://discourse.julialang.org/t/please-read-make-it-easier-to-help-you/14757)

---

<div class="post-metadata">

**Author:** ![dmbates](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmbates/32/44_2.png) [@dmbates](https://discourse.julialang.org/u/dmbates)\
**Post date:** [August 26, 2024, 3:35pm UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/5 "2024-08-26T15:35:58Z")

</div>

Both the R and the GLM.jl implementation of Generalized Linear Models use Iteratively Reweighted Least Squares (IRLS), not EM. The least squares part will fail if the coefficients are undefined due to a singular model matrix (i.e. the X matrix). A general optimizer is less likely to detect the singularity and, depending on the convergence criteria, may well declare convergence near the subspace of possible solutions.

Most analysts, including me, prefer to know when the model is computationally singular. GLM.jl uses a Cholesky factorization of X’WX to solve for the new coefficient vector at each iteration. I have an unregistered package at [GitHub - dmbates/GLMMng.jl: Experimental version of GLM and GLMM fitting in Julia](https://github.com/dmbates/GLMMng.jl) that uses a QR decomposition of √W \* X, which will be slightly slower but better able to handle near singularity, If you want to try that I will add some documentation (right now it is a test-bed with perfunctory documentation). Or, if you could provide a MWE then we can check the condition number of the weighted model matrix.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [October 1, 2024, 2:43pm UTC](https://discourse.julialang.org/t/why-does-glm-jl-fail-to-fit-my-logistic-regression-model-but-scikit-learn-and-r-has-no-issues-i-think-i-found-the-reason/118553/6 "2024-10-01T14:43:54Z")

</div>

This is the same issue I encountered

> <https://github.com/JuliaStats/GLM.jl/issues/426>
>
> \[Car-Training.csv\](https://github.com/JuliaStats/GLM.jl/files/6384056/Car-Traini…ng.csv)
> 
> This is a bit urgent.
> 
> In a home exam that it is ongoing, I have the attached dataset, where students need to estimate models.
> 
> Here are the Julia codes:
> 
> \`\`\`
> using GLM
> using DataFrames
> using CSV
> 
> data = CSV.read( "Car-Training.csv", DataFrame )
> model = @formula( Price ~ Year + Mileage )
> results = lm( model, data )
> \`\`\`
> 
> The output is the following:
> \`\`\`
> StatsModels.TableRegressionModel{LinearModel{GLM.LmResp{Array{Float64,1}},GLM.DensePredChol{Float64,LinearAlgebra.CholeskyPivoted{Float64,Array{Float64,2}}}},Array{Float64,2}}
> 
> Price ~ 1 + Year + Mileage
> 
> Coefficients:
> ─────────────────────────────────────────────────────────────────────────────────
> Coef. Std. Error t Pr(\>|t|) Lower 95% Upper 95%
> ─────────────────────────────────────────────────────────────────────────────────
> (Intercept) 0.0 NaN NaN NaN NaN NaN
> Year 8.17971 0.167978 48.70 \<1e-73 7.84664 8.51278
> Mileage -0.0580528 0.00949846 -6.11 \<1e-7 -0.0768865 -0.0392191
> ─────────────────────────────────────────────────────────────────────────────────
> \`\`\`
> 
> The intercept is not estimated. The other two coefficients are not correct.
> 
> I tried R, SPSS, and Excel. All gave the same results that are different from Julia. I post the results from R below:
> 
> \`\`\`
> R\> data \<- read.csv("Car-Training.csv")
> R\> lm(Price ~ Year + Mileage, data = data )
> 
> Call:
> lm(formula = Price ~ Year + Mileage, data = data)
> 
> R\> summary( lm(Price ~ Year + Mileage, data = data ) )
> 
> Call:
> lm(formula = Price ~ Year + Mileage, data = data)
> 
> Residuals:
> Min 1Q Median 3Q Max 
> -2168 -835 -49 567 5402 
> 
> Coefficients:
> Estimate Std. Error t value Pr(\>|t|)
> (Intercept) -2.19e+06 3.52e+05 -6.21 1.1e-08
> Year 1.10e+03 1.75e+02 6.25 9.0e-09
> Mileage -2.38e-02 9.84e-03 -2.42 0.017
> 
> Residual standard error: 1300 on 104 degrees of freedom
> Multiple R-squared: 0.464, Adjusted R-squared: 0.454 
> F-statistic: 45.1 on 2 and 104 DF, p-value: 8e-15
> \`\`\`
> 
> Did I do something wrong here? Or is there an issue with GLM? If there is an issue with GLM, can a new version be published as soon as possible, so that I can notify the students who are taking this exam right now?
