# How to efficiently find columns of the matrix which are the same?

**URL:** <https://discourse.julialang.org/t/how-to-efficiently-find-columns-of-the-matrix-which-are-the-same/105873>\
**Category:** New to Julia\
**Tags:** question, optimization\
**Created:** [November 6, 2023, 6:56pm UTC](https://discourse.julialang.org/t/how-to-efficiently-find-columns-of-the-matrix-which-are-the-same/105873 "2023-11-06T18:56:47Z")\
**Posts on this page:** 1\
**Showing post:** 11

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [November 7, 2023, 9:39pm UTC](https://discourse.julialang.org/t/how-to-efficiently-find-columns-of-the-matrix-which-are-the-same/105873/11 "2023-11-07T21:39:24Z")

</div>

> [@rocco\_sprmnt21](#):
>
> how does countmap run twice as fast as colmap?

`countmap` [uses `Dict` internals](https://github.com/JuliaStats/StatsBase.jl/blob/77e63df55233d9fc3119bd10144d5b8cc1b9f0ad/src/counts.jl#L286-L296) to avoid hashing each column twice (working around [julia#24454](https://github.com/JuliaLang/julia/issues/24454)), and in this problem hashing the columns is the dominant cost for a sufficiently large matrix.

(Indeed, since hashing is the dominant cost, you could probably speed things up further by wrapping the column views in another type that implements a [custom hash function](https://discourse.julialang.org/t/dictionary-with-custom-hash-function/49168/4). Standard numeric hashing in Julia pays an extra cost so that different types of the same value, e.g. `1.0` vs `1` vs `1.0f0`, yield the same hash. For this application you don’t need that, assuming that all of the matrix elements have the same type, i.e. the `eltype` is concrete. But this kind of voodoo might not be worth the trouble here.)

---

_[View the full topic](https://discourse.julialang.org/t/how-to-efficiently-find-columns-of-the-matrix-which-are-the-same/105873)._
