# Group by criteria

**URL:** <https://discourse.julialang.org/t/group-by-criteria/96246>\
**Category:** General Usage\
**Tags:** iterators\
**Created:** [March 17, 2023, 2:10pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246 "2023-03-17T14:10:03Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 17, 2023, 2:10pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/1 "2023-03-17T14:10:03Z")

</div>

How would you solve the following problem?  
Given the following array of strings we want to find all groups whose group leader starts with “AT”.  
All the elements that are found before the first one starting with “AT” form a group by themselves.

```julia
julia> itr=[randstring("ACTG") for _ in 1:20]
20-element Vector{String}:
 "CTTCCGAG"
 "CCCGTGGT"
 "TCAAGGGT"
 "ATTAGATC"
 "TCTTACAC"
 "TTTCCGCC"
 "TCCGACCG"
 "GTCAGCTA"
 "CATGTTGC"
 "GAGGAACG"
 "GTCAATGC"
 "TACTCATT"
 "ATACTCTA"
 "AATTCACA"
 "AATCATAT"
 "GTATACCT"
 "ATTTTACT"
 "TTCAGAAG"
 "GTTGATGA"
 "GACGGCGG"

julia> mygroupby(itr, criteria)
4-element Vector{Tuple{Vararg{String}}}:
 ("CTTCCGAG", "CCCGTGGT", "TCAAGGGT")
 ("ATTAGATC", "TCTTACAC", "TTTCCGCC", "TCCGACCG", "GTCAGCTA", "CATGTTGC", "GAGGAACG", "GTCAATGC", "TACTCATT")
 ("ATACTCTA", "AATTCACA", "AATCATAT", "GTATACCT")
 ("ATTTTACT", "TTCAGAAG", "GTTGATGA")

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [March 17, 2023, 2:22pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/2 "2023-03-17T14:22:38Z")

</div>

Something like the answers here:

> <https://stackoverflow.com/questions/63738953/julia-idiomatic-way-to-split-vector-to-subvectors-based-on-condition>

with `findall(x -> startswith(x, "AT"), itr)` to find the split points?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [March 17, 2023, 2:27pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/3 "2023-03-17T14:27:04Z")

</div>

Note that if you’re working with DNA/RNA and care about performance you might want to use BioSequences.jl which can be faster than using Strings to store genes.

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [March 17, 2023, 2:55pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/4 "2023-03-17T14:55:31Z")

</div>

```julia
using IterTools

groupby(let c=0; x->c+=startswith(x,"AT") end, itr)

```

gives an iterator returning a group each time:

```julia
julia> groupby(let c=0; x->c+=startswith(x,"AT") end, itr) |> collect
4-element Vector{Vector{String}}:
 ["CTTCCGAG", "CCCGTGGT", "TCAAGGGT"]
 ["ATTAGATC", "TCTTACAC", "TTTCCGCC", "TCCGACCG", "GTCAGCTA", "CATGTTGC", "GAGGAACG", "GTCAATGC", "TACTCATT"]
 ["ATACTCTA", "AATTCACA", "AATCATAT", "GTATACCT"]
 ["ATTTTACT", "TTCAGAAG", "GTTGATGA", "GACGGCGG"]

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 17, 2023, 5:50pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/5 "2023-03-17T17:50:14Z")

</div>

It’s not as good as @Dan’s, but I’m proposing this one too

```julia
reduce((s,c)->startswith(c,"AC") ? push!(s,[c]) : [s[1:end-1];[push!(last(s),c)]],itr, init=[String[]])

```

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [March 17, 2023, 9:25pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/6 "2023-03-17T21:25:49Z")

</div>

This seems to perform well:

```julia
function group2(itr, str)
    ix = [1; findall(x -> startswith(x, str), itr); length(itr)+1]
    return [view(itr, ix[i]:(ix[i+1]-1)) for i in 1:(length(ix)-1)]
end

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 1:56pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/7 "2023-03-18T13:56:56Z")

</div>

this should give the same result

```julia
groupby(let c=true; x->c=startswith(x,"CT") ? !c : c end, itr)|>collect

```

PS  
if group(by,itr) had the analogous form of reduce, for example, with a two-variable function and an init, it would be more flexible.

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [March 18, 2023, 3:40pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/8 "2023-03-18T15:40:49Z")

</div>

Same idea as last one, with some unicode xoring:

```julia
groupby(let c=true; x->c ⊻= startswith(x,"AT") end, itr)|>collect

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 6:06pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/9 "2023-03-18T18:06:16Z")

</div>

a variant

```julia
reduce((d,c)->startswith(c,"AC") ? (d[1]+=1;d[2][d[1]]=[c];d) : (push!(get!(()->[],d[2],d[1]),c);d),itr, init=[0,Dict{Int, Vector{String}}()])

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 18, 2023, 6:12pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/10 "2023-03-18T18:12:40Z")

</div>

This answer seems pretty convoluted. Is it really helping answer the question? Julia obviously has a lot of ways to solve every problem, but I don’t think its useful to outline every single one.

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 6:20pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/11 "2023-03-18T18:20:24Z")

</div>

Yes. I know that the ones I wrote are convoluted and not very useful for the purpose of the problem.  
But they are “useful” exercises… at least

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 8:37pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/12 "2023-03-18T20:37:30Z")

</div>

Now supposing we have the following list of strings, we want to group them (scrolling them from first to last) in such a way that the strings with the “not too different root” are in the same group.

```julia
20-element Vector{String}:
 "GCCCCACA"
 "TCCATTTT"
 "TGCCCCGT"
 "GGGTTAGT"
 "GATTGAAG"
 "CCCGTAAT"
 "GGACCACT"
 "GGCGACAC"
 "GCTCTGTA"
 "AGTCCTGC"
 "TCATTACC"
 "GAAGTAGT"
 "AACTCGAA"
 "ATGACCCT"
 "TATTATCG"
 "CCTTGTCA"
 "TGCTCGCC"
 "CCGCATGG"
 "CTCTTTTG"
 "TTATGGAC"

```

more precisely with `dist(n1,n2)>0` where

`dist(n1,n2)= length(intersect(first(n1,3),first(n2,3)))`

something like this should come

```julia
"GCC" => ["GCCCCACA", "TCCATTTT", "TGCCCCGT", "GGGTTAGT", "GATTGAAG", "CCCGTAAT", "GGACCACT", "GGCGACAC", "GCTCTGTA", "AGTCCTGC", "TCATTACC", "GAAGTAGT", "AACTCGAA", "ATGACCCT"], 
"TAT" => ["TATTATCG", "CCTTGTCA", "TGCTCGCC"], 
"CCG" => ["CCGCATGG", "CTCTTTTG"], 
"TTA" => ["TTATGGAC"]

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 18, 2023, 8:45pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/13 "2023-03-18T20:45:55Z")

</div>

This problem is poorly specified. What about ties? What about transitive differences?

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 8:51pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/14 "2023-03-18T20:51:46Z")

</div>

If I understand the meaning of the observation, I specify better what I have already written  
The groups are made up of consecutive elements of the list, similar to what the group function of itertools does

```julia
# Group consecutive values that share the same result of applying f.

```

---

<div class="post-metadata">

**Author:** ![pdeffebach](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/pdeffebach/32/10320_2.png) [@pdeffebach](https://discourse.julialang.org/u/pdeffebach)\
**Post date:** [March 18, 2023, 9:02pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/15 "2023-03-18T21:02:48Z")

</div>

The following should work

```julia
julia> using SplitApplyCombine;

julia> itr = [randstring("ACTG") for _ in 1:20];

julia> dist(n1, n2)= length(intersect(first(n1, 3), first(n2, 3)));

julia> diffs = similar(itr, Bool);

julia> for i in eachindex(itr)
           if i == firstindex(itr)
               diffs[i] = false
           else
               diffs[i] = dist(diffs[i-1], diffs[i]) > 0
           end
       end;

julia> diffs_groups = cumsum(diffs);

julia> group(diffs_goups, itr)
12-element Dictionaries.Dictionary{Int64, Vector{String}}
  0 │ ["GCGGGCCG"]
  1 │ ["ATGTCTGT", "TGTTATCA"]
  2 │ ["GCAAGAAC"]
  3 │ ["AGTTCCTA"]
  4 │ ["GGCGTACC", "GCTATGAT"]
  5 │ ["GACTCACG", "CAATGGTG"]
  6 │ ["AGAGCCTG"]
  7 │ ["CACGGGGA", "CATGATTG", "TTCTCAAG", "GAGTCTAT"]
  8 │ ["CCATGATA", "TCGCCGTA"]
  9 │ ["GGCATATT"]
 10 │ ["CTCTAACT", "CGTACGCG"]
 11 │ ["CTCGCTGT"]

```

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 18, 2023, 10:23pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/16 "2023-03-18T22:23:50Z")

</div>

in this case, the whole list is one group.  
For each string, in the first 3 letters there is at least one of the first 3 of the list leader

```julia
1-element Dictionary{String, Vector{String}}
 "GCG" │ ["GCGGGCCG", "ATGTCTGT", "TGTTATCA", "GCAAGAAC", "AGTTCCTA", "GGCGTACC…

```

another example to better clarify the goal

```julia
julia> itr=[randstring("ACTG") for _ in 1:20]
20-element Vector{String}:
 "GCTATAAA"
 "CCTTCGGG"
 "TACGCGAG"
 "AAAAACGA"
 "ACCCGCTT"
 "GTGCGATC"
 "ATCTCGCG"
 "GGATATCA"
 "AACATATG"
 "GGTGCAGC"
 "GTTATGGC"
 "CGAGCTAA"
 "CCCTGCCA"
 "AAGACATC"
 "AGTATCCA"
 "CGATCGGT"
 "AGCCACAC"
 "TTCTATCA"
 "GCTTAGTC"
 "AGCATACA"

julia> 

julia> ...
8-element Vector{Vector{String}}:
 ["GCTATAAA", "CCTTCGGG", "TACGCGAG"]
 ["AAAAACGA", "ACCCGCTT"]
 ["GTGCGATC", "ATCTCGCG", "GGATATCA"]
 ["AACATATG"]
 ["GGTGCAGC", "GTTATGGC", "CGAGCTAA"]
 ["CCCTGCCA"]
 ["AAGACATC", "AGTATCCA", "CGATCGGT", "AGCCACAC"]
 ["TTCTATCA", "GCTTAGTC", "AGCATACA"]

```

from the following it is better seen that each group leader has no letters in common with the previous group leader

```julia
8-element Dictionary{String, Vector{String}}
 "GCT" │ ["GCTATAAA", "CCTTCGGG", "TACGCGAG"]
 "AAA" │ ["AAAAACGA", "ACCCGCTT"]
 "GTG" │ ["GTGCGATC", "ATCTCGCG", "GGATATCA"]
 "AAC" │ ["AACATATG"]
 "GGT" │ ["GGTGCAGC", "GTTATGGC", "CGAGCTAA"]
 "CCC" │ ["CCCTGCCA"]
 "AAG" │ ["AAGACATC", "AGTATCCA", "CGATCGGT", "AGCCACAC"]
 "TTC" │ ["TTCTATCA", "GCTTAGTC", "AGCATACA"]

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [March 18, 2023, 11:40pm UTC](https://discourse.julialang.org/t/group-by-criteria/96246/17 "2023-03-18T23:40:27Z")

</div>

```julia
julia> using IterTools

julia> groupby(let c=""; x->c = any(∈(c),first(x,3)) ? c : first(x,3) end, itr)|>collect
8-element Vector{Vector{String}}:
 ["GCTATAAA", "CCTTCGGG", "TACGCGAG"]
 ["AAAAACGA", "ACCCGCTT"]
 ["GTGCGATC", "ATCTCGCG", "GGATATCA"]
 ["AACATATG"]
 ["GGTGCAGC", "GTTATGGC", "CGAGCTAA"]
 ["CCCTGCCA"]
 ["AAGACATC", "AGTATCCA", "CGATCGGT", "AGCCACAC"]
 ["TTCTATCA", "GCTTAGTC", "AGCATACA"]

```

This works. What is the motivation behind these exercises?

---

<div class="post-metadata">

**Author:** ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)\
**Post date:** [March 19, 2023, 8:10am UTC](https://discourse.julialang.org/t/group-by-criteria/96246/18 "2023-03-19T08:10:38Z")

</div>

I had a vague memory of [this](https://github.com/JuliaData/SplitApplyCombine.jl/pull/38) that I have now found, and I wanted to check if the use you showed of groupby() responded to some of the examples proposed there.

this last example is because I wanted to find a case where one of those convoluted propositions with the use of reducers and dictionaries would be useful, thinking that groupby can’t deal with such cases

```julia
reduce((d,c)->dist(c,d[1])<1 ? (d[1]=first(c,3);insert!(d[2], d[1], [c]);d) : (push!(get!(()->[],d[2],d[1]),c);d),itr, init=["",Dictionary{String, Vector{String}}()])[2]

```

but then, thinking about it better, I found the way (similar to what you proposed) to use groupby also in this case

```julia

dist(n1,n2)= length(intersect(first(n1,3),first(n2,3)))
groupby(let c=first(itr[1],3); x->dist(x,c)==0 ? c=first(x,3) : c end, itr)|>collect

# or

dist1(n1,n2)=isempty(intersect(first(n1,3),first(n2,3))) ? first(n2,3) : missing

groupby(let c=""; x->c=coalesce(dist1(c,x),c) end, itr)|>collect

```
