# How to remove specific collumn from ModelFrame (StatsModels.jl)?

**URL:** <https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864>\
**Category:** Statistics\
**Created:** [November 28, 2023, 8:12pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864 "2023-11-28T20:12:19Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [November 28, 2023, 8:12pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/1 "2023-11-28T20:12:19Z")

</div>

Hello!

I have ModelFrame from data:

```julia
mf = ModelFrame(@formula(y ~ 1 + x*y), df))

```

where ‘x’ and ‘y’ - categorical. But I have rank-deficient `modelmatrix(mf)` and want to drop rank-deficient columns, I can identify them, but have no idea how to rebuild ModelFrame with right coefnames ets? Any idea how to do that?

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [November 28, 2023, 8:59pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/2 "2023-11-28T20:59:58Z")

</div>

Hi!  
Sharing from experience: If you give a runnable example (MWE) and some sort of indication of desired output for that example - the answers are so much quicker.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 5, 2023, 11:58am UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/3 "2023-12-05T11:58:20Z")

</div>

Hi!  
I need to make `ModelFrame` that give me full-rank design matrix.

For example I have data:

```julia
using StatsModels, LinearAlgebra

var = rand(10)
fe1 = ["a","a","a","a","b","b","b","b","b","b"] 
fe2 = ["1","1","2","2","1","1","2","2","2","2"]
df = DataFrame(var = var, fe1 = fe1, fe2 = fe2)

f = @formula(var ~ fe1 & fe2)
s = schema(f, df)
as = apply_schema(f, s, LinearModel)
mf = ModelFrame(as, s, df, LinearModel)
mm = modelmatrix(mf) # < This is rank-deficient
cn = coefnames(mf)

# I can find all linear independent columns:
qrd = qr(mm)
cols = findall(x-> abs(x) > 1e-8, diag(qrd.R))

# and i can get good design matrix for modelling and coefnames vector
fr_mm = mm[:, cols]
fr_cn = cn[cols]

# do something with mm (ModelFrame)
# ???

fr_mm == modelmatrix(mf) # should be true

```

Then I want to modify `mf` (`ModelFrame`) and remove linear dependent columns (5) to get good modelmatrix (`mm = modelmatrix(mf)`) any time further, **how to do that**?? - this is main question.

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [December 5, 2023, 12:56pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/4 "2023-12-05T12:56:22Z")

</div>

> [@PharmCat](#):
>
> ```julia
> cols = findall(x-> abs(x) > 1e-8, diag(qrd.R))
> 
> # and i can get good design matrix for modelling and coefnames vector
> fr_mm = mm[:, cols]
> 
> ```

I think the columns in the QR factorization are permuted from the factorized matrix, and so indexing the `mm` matrix with `cols` doesn’t give anything useful.

The problem with finding a set of independent columns is that there are many sets of independent columns and it isn’t clear which one to choose.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 5, 2023, 2:17pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/5 "2023-12-05T14:17:56Z")

</div>

You can find collinear columns with QR decomposition:“[The Behavior of the QR-Factorization Algorithm with Column Pivoting](http://pages.stat.wisc.edu/%7Ebwu62/771/engler1997.pdf)” by Engler (1997)).

In this example I can improve:

```julia
qrd = qr(mm' * mm)
cols = findall(x-> abs(x) > 1e-8, diag(qrd.R))

```

My question not about how to find collinear columns (I think it can be done not only with QR, with SVD or spectral decomposition), I ask about how to drop certain column from ModelFrame?

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 5, 2023, 9:39pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/6 "2023-12-05T21:39:58Z")

</div>

`ModelFrame` is [deprecated](https://juliastats.org/StatsModels.jl/latest/formula/):

> The `ModelFrame` and `ModelMatrix` types can still be used to do this transformation, but this is only to preserve some backwards compatibility. Package authors who would like to include support for fitting models from a `@formula` are **strongly** encouraged to directly use `schema`, `apply_schema`, and `modelcols` to handle the table-to-matrix transformations they need.

Why do you need it? Instead you could call `modelcols(as, df)` and `coefnames(as)`. This is what GLM.jl does on git master for example.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 5, 2023, 11:36pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/7 "2023-12-05T23:36:08Z")

</div>

Hi! Thank you! So… Question can be changed: what can i do with `FormulaTerm` to get full-rank matrix further:

```julia

var = rand(10)
fe1 = ["a","a","a","a","b","b","b","b","b","b"] 
fe2 = ["1","1","2","2","1","1","2","2","2","2"]
df = DataFrame(var = var, fe1 = fe1, fe2 = fe2)

f = @formula(var ~ fe1 & fe2)
s = schema(f, df)
as = apply_schema(f, s, LinearModel)
Y,mm = modelcols(as, df) 
rn,cn = coefnames(as)

qrd = qr(mm' * mm)
cols = findall(x-> abs(x) > 1e-8, diag(qrd.R))

fr_mm = mm[:, cols]
fr_cn = cn[cols]

# do something with as (FormulaTerm)
# ???
# some code

Y,mm = modelcols(as, df)
fr_mm == mm # true

```

P.S. Suppose, I support ModelFrame because I use `assign` from `ModelMatrix`

```julia
mf = ModelFrame(f, sch, data, MetidaModel)
mm = ModelMatrix(mf)
# then somewhere:
mm.assign

```

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 6, 2023, 8:45am UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/8 "2023-12-06T08:45:04Z")

</div>

Ah so what you want is to be able to compute the equivalent of `mm.assign`? Then I think you can do the same as this (`t` being `as.rhs`):

> <https://github.com/JuliaStats/StatsModels.jl/blob/df862a601947de3912d190cb2f1d9e1d751ad4da/src/modelframe.jl#L219-L222>

And then you can drop the entries corresponding to the columns you dropped.

But I don’t think you can nor need to modify the formula itself.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 6, 2023, 11:18am UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/9 "2023-12-06T11:18:51Z")

</div>

Hi! Thank you very much for advice about `assign`, but it is not main question) I still ask about

> how to modify FormulaTerm to get full-rank design matrix

.

So why I need exactly that:

I am making package for statistical models and it can works only with full-rank design matrices (I try to made tweaks to avoid problems, but it not works or works very unstable. and only one appropriate way - to use full-rank matrix).

So… I use StatModels and now store ModelFrame in model object for compatibility with some packages like Effects.jl  
Now I take design matrix (X) with StatsModels from ModelFrame (and change that how you advice) and work with it, but if I need full-rank design matrix - I should modify it, and after that it no more corresponds to object I saved (ModelFrame or FormulaTerm) and if other package try to get coefficients from `FormulaTerm` it would not match coefficients estimated from model.

 ![image](https://global.discourse-cdn.com/julialang/original/3X/e/0/e09d94ef916b080f05341b0068a1eb1e0960b2df.png)

So, for correct working I need to modify FormulaTerm to make it structure corresponsable to real coefficients and covariance.  
Other way - to drop support of StatsModels, but in this way I lost ability to use Effects.jl  
Now I think that working around StatsModels is still appropriate way for package developing… thats why I try to get way to modify FormulaTerm to get full-rank design matrix

---

<div class="post-metadata">

**Author:** ![nalimilan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nalimilan/32/147_2.png) [@nalimilan](https://discourse.julialang.org/u/nalimilan)\
**Post date:** [December 10, 2023, 4:39pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/10 "2023-12-10T16:39:44Z")

</div>

OK. Unfortunately that doesn’t seem easy, in particular because it implies the interaction between multiple terms rather than a single term. @dave.f.kleinschmidt has an idea.

A possible hack would be to fit the model using the modified matrix with collinear columns dropped, but still return the full coefficients and covariance matrix with `NaN` for omitted columns.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 12, 2023, 3:36pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/11 "2023-12-12T15:36:37Z")

</div>

Thank you! I’ll try to come up with something.

---

<div class="post-metadata">

**Author:** ![PharmCat](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pharmcat/32/6953_2.png) [@PharmCat](https://discourse.julialang.org/u/PharmCat)\
**Post date:** [December 24, 2023, 2:32pm UTC](https://discourse.julialang.org/t/how-to-remove-specific-collumn-from-modelframe-statsmodels-jl/106864/12 "2023-12-24T14:32:55Z")

</div>

Hi!

I add this to my Proj:

```julia
asgn(f::FormulaTerm) = asgn(f.rhs)
asgn(t) = mapreduce(((i,t), ) -> i*ones(StatsModels.width(t)),
                    append!,
                    enumerate(StatsModels.vectorize(t)),
                    init=Int[])

```

so I have:

```julia
 Proj.asgn(f)
6-element Vector{Int64}:
 1
 1
 1
 1
 1
 1

```

and

```julia
julia> StatsModels.asgn(f)
6-element Vector{Int64}:
 2
 2
 3
 3
 3
 4

```

Is any thoughts how to fix it?
