# Remove identical columns from matrix

**URL:** <https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378>\
**Category:** Statistics\
**Created:** [June 4, 2021, 6:47am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378 "2021-06-04T06:47:24Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![amrods](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amrods/32/2543_2.png) [@amrods](https://discourse.julialang.org/u/amrods)\
**Post date:** [June 4, 2021, 6:47am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/1 "2021-06-04T06:47:24Z")

</div>

Suppose I have a matrix with several columns. What’s an easy way to remove identical columns? To provide an example, I want to get matrix `x` below if I start from matrix `xx`.

```julia
x = rand(10, 5)
xx = [x x[:, 2] x[:, 5]]

```

---

<div class="post-metadata">

**Author:** ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)\
**Post date:** [June 4, 2021, 6:56am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/2 "2021-06-04T06:56:57Z")

</div>

```julia
hcat(unique(eachcol(xx))...)

```

---

<div class="post-metadata">

**Author:** ![amrods](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amrods/32/2543_2.png) [@amrods](https://discourse.julialang.org/u/amrods)\
**Post date:** [June 4, 2021, 7:13am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/3 "2021-06-04T07:13:05Z")

</div>

I was expecting some kind of algorithm. There is always a better way! 🙂

---

<div class="post-metadata">

**Author:** ![amrods](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amrods/32/2543_2.png) [@amrods](https://discourse.julialang.org/u/amrods)\
**Post date:** [June 4, 2021, 7:21am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/4 "2021-06-04T07:21:12Z")

</div>

Challenge question: what if I wanted to eliminate collinear columns, rather than identical columns?

---

<div class="post-metadata">

**Author:** ![ettersi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ettersi/32/6829_2.png) [@ettersi](https://discourse.julialang.org/u/ettersi)\
**Post date:** [June 4, 2021, 7:28am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/5 "2021-06-04T07:28:00Z")

</div>

`hcat(unique(normalize.(eachcol(xx)))...)` should do the trick

---

<div class="post-metadata">

**Author:** ![amrods](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amrods/32/2543_2.png) [@amrods](https://discourse.julialang.org/u/amrods)\
**Post date:** [June 4, 2021, 7:51am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/6 "2021-06-04T07:51:28Z")

</div>

`normalize`?

---

<div class="post-metadata">

**Author:** ![ettersi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ettersi/32/6829_2.png) [@ettersi](https://discourse.julialang.org/u/ettersi)\
**Post date:** [June 4, 2021, 7:52am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/7 "2021-06-04T07:52:06Z")

</div>

It’s provided by the `LinearAlgebra` package.

---

<div class="post-metadata">

**Author:** ![amrods](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/amrods/32/2543_2.png) [@amrods](https://discourse.julialang.org/u/amrods)\
**Post date:** [June 4, 2021, 7:56am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/8 "2021-06-04T07:56:59Z")

</div>

That returns a transform of matrix `x`, which I don’t want, and still returns a 10x6 matrix rather than `x`.

```julia
x = rand(10, 5)
xx = [x x[:, 2] x[:, 5] x[:, 1] .+ 2 .* x[:, 3]]
hcat(unique(normalize.(eachcol(xx)))...)

```

---

<div class="post-metadata">

**Author:** ![ettersi](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ettersi/32/6829_2.png) [@ettersi](https://discourse.julialang.org/u/ettersi)\
**Post date:** [June 4, 2021, 8:04am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/9 "2021-06-04T08:04:29Z")

</div>

> [@amrods](#):
>
> still returns a 10x6 matrix rather than `x` .

That’s because `x[:, 1] .+ 2 .* x[:, 3]` is a linear combination of the columns of `x` but not collinear with any column in `x`.

---

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [June 4, 2021, 8:16am UTC](https://discourse.julialang.org/t/remove-identical-columns-from-matrix/62378/10 "2021-06-04T08:16:44Z")

</div>

If you want to get a maximal linearly independent subset of the original columns (in other words, find columns of A which build a basis of the column space) you can do:

```julia
using RowEchelon

_, pivots = rref_with_pivots(A)
A[:, pivots]

```

Or you can use `svd(A).U[:, 1:rank(A)]` if you just want a basis of the column space, and don’t care if the basis vectors are columns of `A`. (Note that `rank` does its own SVD so if performance is an issue, get the singular values from the first `svd` call and check yourself which are close to 0, which is what `rank` does.)
