# Logistic Regression Problem

**URL:** <https://discourse.julialang.org/t/logistic-regression-problem/5417>\
**Category:** Statistics\
**Created:** [August 16, 2017, 7:37pm UTC](https://discourse.julialang.org/t/logistic-regression-problem/5417 "2017-08-16T19:37:45Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Antonio\_Loureiro](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antonio_loureiro/32/15257_2.png) [@Antonio\_Loureiro](https://discourse.julialang.org/u/Antonio_Loureiro)\
**Post date:** [August 16, 2017, 7:37pm UTC](https://discourse.julialang.org/t/logistic-regression-problem/5417/1 "2017-08-16T19:37:45Z")

</div>

I’m trying to use a Logistic Regression algorithm to find a classification model, but i get stuck with an error “failure to converge after 30 iterations” i have changed the maxIter arg to an higher value, but the error only disappears with a very high value of iterations, in the small example below with a population of 1000 elements and with a very simple implicit model i need 2000 iterations! Am i doing something wrong?

> Blockquote  
> df=DataFrame()  
> n=1000  
> df[:x]=rand(n)  
> df[:y]=rand(n)  
> df[:z]=rand(n)  
> df[:valid]=map((x,y)-\>(x_2-y_6)\>0? true : false,df[:x],df[:y])  
> glm(@formula(valid ~ x+y+z), df, Binomial(), LogitLink())

---

<div class="post-metadata">

**Author:** ![andreasnoack](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/andreasnoack/32/27_2.png) [@andreasnoack](https://discourse.julialang.org/u/andreasnoack)\
**Post date:** [August 16, 2017, 7:58pm UTC](https://discourse.julialang.org/t/logistic-regression-problem/5417/2 "2017-08-16T19:58:10Z")

</div>

Since there is no noise added to the linear predictor, it predicts the outcomes perfectly and, as a consequence, the MLE doesn’t exist, i.e. the likelihood function doesn’t have an optimum but keeps growing as one or more of the model parameters diverge. See e.g. [FAQ What is complete or quasi-complete separation in logistic/probit regression and how do we deal with them?](https://stats.idre.ucla.edu/other/mult-pkg/faq/general/faqwhat-is-complete-or-quasi-complete-separation-in-logisticprobit-regression-and-how-do-we-deal-with-them/). In ML, people usually add some regularization to ensure that an optimum exists but GLM doesn’t add regularization.

---

<div class="post-metadata">

**Author:** ![Antonio\_Loureiro](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antonio_loureiro/32/15257_2.png) [@Antonio\_Loureiro](https://discourse.julialang.org/u/Antonio_Loureiro)\
**Post date:** [August 16, 2017, 9:57pm UTC](https://discourse.julialang.org/t/logistic-regression-problem/5417/3 "2017-08-16T21:57:03Z")

</div>

Thks, that was a perfect answer!

---

<div class="post-metadata">

**Author:** ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)\
**Post date:** [August 16, 2017, 9:59pm UTC](https://discourse.julialang.org/t/logistic-regression-problem/5417/4 "2017-08-16T21:59:27Z")

</div>

Is it planned functionality to add some simple normalization to GLM? If, as I’m assuming, the fitting of a GLM is essentially Newton-Raphson algorithm, it should be easy to add some optional L2 cost parameter. It could be a big plus for usability, otherwise it can be confusing for a new user that, as soon as your problem is a bit ill-defined/unstable, you need to switch to I guess [GLMnet](https://github.com/JuliaStats/GLMNet.jl) or [Lasso](https://github.com/simonster/Lasso.jl) which is a different package with different syntax etc.
