# Finding sub-arrays in an array?

**URL:** https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700
**Category:** General Usage
**Tags:** data, chemistry
**Created:** [April 13, 2026, 4:50pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700 "2026-04-13T16:50:53Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![sophibug](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sophibug/32/210241_2.png) [@sophibug](https://discourse.julialang.org/u/sophibug)
#### Post date: [April 13, 2026, 4:50pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/1 "2026-04-13T16:50:53Z")

</div>

I’m putting this here because this is a chemistry usage but it’s not really a chemistry-specific case.

I have an array that has a bunch of states formatted like this:

> 1.0 1.0 0.0 1.0 0.0 1.0 6380.35  
> 1.0 1.0 1.0 1.0 0.0 1.0 6316.91  
> 2.0 1.0 1.0 2.0 0.0 2.0 6444.27  
> 2.0 1.0 2.0 1.0 0.0 1.0 11086.5

For any given line, 1:3 is a description of one state, 4:6 is the description of a second state, and 7 is the frequency of the difference.

I have created a list of all the states (which is all the values 1:3 and 4:6 listed in an nx3 array) and of all the unique states, and I need to identify which states only appear once in the list. However,

> findall(isequal(unst[1,:]), states[1:3,:])

Is not working – it provides an empty Cartesian index

> CartesianIndex{2}

Preferably, I would be able to find the index of every unique state in the overall nx7 list. Does anyone know how to do this?

Thanks!

---

<div class="post-metadata">

### Author: ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)
#### Post date: [April 13, 2026, 5:07pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/2 "2026-04-13T17:07:13Z")

</div>

The issue is that `isequal` operates on the entire array, rather than columnwise like you seem to want. SInce the entire matrix does not match `unst[1,:]`, you get a single `false` as a result.

Does

```julia
findall(isequal(unst[1,:]), eachcol(@view(states[1:3, :]))) # @view is optional

```

do what you want? This should slice the matrix into each 3-tall column and compare each of those to `unst[1, :]`.

For example,

```julia-repl
julia> findall(isequal([1;2;3]), eachcol([1;1;1;; 1;2;3;; 3;2;1;; 2;2;2;;]))
1-element Vector{Int64}:
 2

```

---

<div class="post-metadata">

### Author: ![sophibug](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sophibug/32/210241_2.png) [@sophibug](https://discourse.julialang.org/u/sophibug)
#### Post date: [April 13, 2026, 7:58pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/3 "2026-04-13T19:58:34Z")

</div>

I’m actually trying to match rows, not columns. States is formatted like

```julia-auto
  1 1 0
  1 1 1
  2 1 1
  2 1 2
  2 2 0
  2 2 1

```

and unst is formatted like

```julia-auto
  1 1 0
  1 1 1
  2 1 1
  2 1 2
  2 2 0

```

---

<div class="post-metadata">

### Author: ![mikmoore](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mikmoore/32/31109_2.png) [@mikmoore](https://discourse.julialang.org/u/mikmoore)
#### Post date: [April 13, 2026, 8:25pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/4 "2026-04-13T20:25:29Z")

</div>

> [@sophibug](#):
>
> I’m actually trying to match rows, not columns. States is formatted like

The rest of my earlier comment is likely still relevant except that you might need to use `eachrow(states[:, 1:3])` instead of `eachcol(states[1:3, :])`.

* * *

In general, the NumPy/MATLAB/etc pattern of smooshing all your data into a mega-array and then slicing-and-dicing it during processing is not necessary in Julia (or other languages where non-arrays are performant). Where relevant, you might consider using types to organize your data a little more carefully. For example:

```julia-repl
julia> struct StateDescriptor
           state1::NTuple{3, Int} # maybe even make a special type for these
           state2::NTuple{3, Int}
           freq_diff::Float64
       end

julia> bunch_of_states = [StateDescriptor((1,1,0), (1,0,1), 6380.35), StateDescriptor((1,1,1), (1,0,1), 6316.91), StateDescriptor((2,1,1), (2,0,2), 6444.27), StateDescriptor((2,1,2), (1,0,1), 11086.5)]
4-element Vector{StateDescriptor}:
 StateDescriptor((1, 1, 0), (1, 0, 1), 6380.35)
 StateDescriptor((1, 1, 1), (1, 0, 1), 6316.91)
 StateDescriptor((2, 1, 1), (2, 0, 2), 6444.27)
 StateDescriptor((2, 1, 2), (1, 0, 1), 11086.5)

julia> all_states = vcat([x.state1 for x in bunch_of_states], [x.state2 for x in bunch_of_states])
8-element Vector{Tuple{Int64, Int64, Int64}}:
 (1, 1, 0)
 (1, 1, 1)
 (2, 1, 1)
 (2, 1, 2)
 (1, 0, 1)
 (1, 0, 1)
 (2, 0, 2)
 (1, 0, 1)

julia> unique_states = unique(all_states)
6-element Vector{Tuple{Int64, Int64, Int64}}:
 (1, 1, 0)
 (1, 1, 1)
 (2, 1, 1)
 (2, 1, 2)
 (1, 0, 1)
 (2, 0, 2)

julia> findall(isequal(unique_states[1]), x.state1 for x in bunch_of_states) # (1, 1, 0) in state1
1-element Vector{Int64}:
 1

julia> findall(isequal(unique_states[5]), x.state2 for x in bunch_of_states) # (1, 0, 1) in state2
3-element Vector{Int64}:
 1
 2
 4

```

Done this way, there isn’t even a question about rows versus columns. I’m sure you’d want to make some adjustments to what I’ve suggested based on your full use case, but in general I find data/code like this much easier to reason about.

---

<div class="post-metadata">

### Author: ![sophibug](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/sophibug/32/210241_2.png) [@sophibug](https://discourse.julialang.org/u/sophibug)
#### Post date: [April 13, 2026, 9:23pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/5 "2026-04-13T21:23:21Z")

</div>

Interesting! I have it in an array like this because it’s read in from a file, but I could do some data manipulation like that before I got to this point.

Thanks for your help; that worked.

---

<div class="post-metadata">

### Author: ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)
#### Post date: [April 14, 2026, 12:05am UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/6 "2026-04-14T00:05:35Z")

</div>

If the intention was to view the list as a list of possible transitions between states, and furthermore to find the states which appear once and only once in the transition list, then maybe the following can help:

```julia-auto
# run the code in mikmoore's post...
using StatsBase

cm = countmap(all_states)
# Dict{Tuple{Int64, Int64, Int64}, Int64} with 6 entries:
# (1, 1, 1) => 1
# (1, 0, 1) => 3
# ...

only_once_state_indices = # indices inside `all_states`
  [i for (i,s) in enumerate(all_states) if last(cm[s])==1]
# 5-element Vector{Int64}:
# 1
# 2
# ...

only_once_state_transitions = # indices in transition list
  [i for (i,s) in enumerate(bunch_of_states) 
    if ( last(cm[s.state1])==1 || last(cm[s.state2])==1 ) ]
# 4-element Vector{Int64}:
# 1
# 2
# 3
# 4
# apparantly, each transition has one state which appears only once

```

Also, it would be better to rename `bunch_of_states` to `bunch_of_transitions`.

---

<div class="post-metadata">

### Author: ![kapple](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kapple/32/218915_2.png) [@kapple](https://discourse.julialang.org/u/kapple)
#### Post date: [April 14, 2026, 5:56am UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/7 "2026-04-14T05:56:50Z")

</div>

As an alternative to using `structs` you might find it useful to load data to a `DimStack` of `DimArray`s from DimensionalData.jl, depending on what other things you will be doing with your data.

---

<div class="post-metadata">

### Author: ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)
#### Post date: [April 14, 2026, 3:50pm UTC](https://discourse.julialang.org/t/finding-sub-arrays-in-an-array/136700/8 "2026-04-14T15:50:02Z")

</div>

> [@sophibug](#):
>
> Interesting! I have it in an array like this because it’s read in from a file, but I could do some data manipulation like that before I got to this point.

Assuming the file is in any of the common formats, try reading it with the corresponding julia package – eg, both delimited and fixed-width textual formats are well-supported in Julia.

Then, you will immediately have 1d-array-of-namedtuples instead of 2d-array, and these are very convenient and efficient to manipulation in Julia. Defining custom structs or using fancy array wrappers can be useful, but those are generally further potential steps – you already get great usability and performance from just namedtuples all the way 🙂
