# Multi column indexing

**URL:** <https://discourse.julialang.org/t/multi-column-indexing/66265>\
**Category:** New to Julia\
**Tags:** arrays\
**Created:** [August 12, 2021, 12:29pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265 "2021-08-12T12:29:37Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jpmorr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpmorr/32/26819_2.png) [@jpmorr](https://discourse.julialang.org/u/jpmorr)\
**Post date:** [August 12, 2021, 12:29pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265/1 "2021-08-12T12:29:37Z")

</div>

I’m still new to Julia but there’s still some basics that I’m still clearly not understanding. Take the following example which has confused me all morning.

> storage\_indx\_mult = hcat([12, 165, 55, 66, 89, 101, 2], [68, 135, 409, 222,6, 818,7])  
> column\_index\_mult = [1,1,2,2,1,1,2]  
> id\_index = [6,4,5,2,1,7,3]  
> test\_ids\_mult = storage\_indx\_mult[id\_index, column\_index\_mult]

In this case I’m expecting to get back a single vector:

> test\_ids\_mult\_expected = [101, 66, 6, 135, 12, 2, 409]

But Julia gives back a full matrix:

> test\_ids\_mult = storage\_indx\_mult[id\_index, column\_index\_mult]  
> 7×7 Matrix{Int64}:  
> 101 101 818 818 101 101 818  
> 66 66 222 222 66 66 222  
> 89 89 6 6 89 89 6  
> 165 165 135 135 165 165 135  
> 12 12 66 66 12 12 66  
> 2 2 7 7 2 2 7  
> 55 55 409 409 55 55 409

As I said, I’m clearly not understanding something about Julia and Indexing. I think coming from a python background is confusing my thinking. How do I extract the single vector I want from the two index vectors?

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [August 12, 2021, 12:47pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265/2 "2021-08-12T12:47:52Z")

</div>

You can try this:

```julia
getindex.(Ref(storage_indx_mult), id_index, column_index_mult)

```

---

<div class="post-metadata">

**Author:** ![jpmorr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpmorr/32/26819_2.png) [@jpmorr](https://discourse.julialang.org/u/jpmorr)\
**Post date:** [August 12, 2021, 1:19pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265/3 "2021-08-12T13:19:56Z")

</div>

> [@rafael.guerra](#):
>
> `getindex.(Ref(storage_indx_mult), id_index, column_index_mult)`

Thanks. I would never have figured this out. Seems rather unintuitive compared to what I’m used to in python.

---

<div class="post-metadata">

**Author:** ![jipolanco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jipolanco/32/12129_2.png) [@jipolanco](https://discourse.julialang.org/u/jipolanco)\
**Post date:** [August 12, 2021, 1:28pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265/4 "2021-08-12T13:28:40Z")

</div>

This alternative may seem more intuitive (and closer to python):

```julia
[storage_indx_mult[i, j] for (i, j) in zip(id_index, column_index_mult)]

```

Note that, in Julia, your `storage_indx_mult[id_index, column_index_mult]` is actually equivalent to `[storage_indx_mult[i, j] for i in id_index, j in column_index_mult]`.

Yet another option:

```julia
storage_indx_mult[CartesianIndex.(id_index, column_index_mult)]

```

---

<div class="post-metadata">

**Author:** ![jpmorr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpmorr/32/26819_2.png) [@jpmorr](https://discourse.julialang.org/u/jpmorr)\
**Post date:** [August 12, 2021, 1:41pm UTC](https://discourse.julialang.org/t/multi-column-indexing/66265/5 "2021-08-12T13:41:54Z")

</div>

> [@jipolanco](#):
>
> `storage_indx_mult[CartesianIndex.(id_index, column_index_mult)]`

Thanks for these other options. The cartesian index seems the most straightforward as list comprehensions are always easy to mess up I find!

I did a quick check with `BencmarkTools` to check perfromance and for a typical dataset I would be working on with about 100K rows, `getindex` seems to be the quickest.

> @btime pIDs[CartesianIndex.(ids[:,1], cc2[:, 1].+1)]  
> 216.600 μs (14 allocations: 1.78 MiB)  
> @btime [pIDs[i, j] for (i, j) in zip(ids[:,1], cc2[:, 1].+1)]  
> 1.886 ms (99894 allocations: 2.91 MiB)  
> @btime getindex.(Ref(pIDs), ids[:,1], cc2[:, 1].+1)  
> 107.900 μs (13 allocations: 1010.94 KiB)
