# PartitionBy, retaining key

**URL:** <https://discourse.julialang.org/t/partitionby-retaining-key/62820>\
**Category:** General Usage\
**Tags:** transducers\
**Created:** [June 13, 2021, 1:18am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820 "2021-06-13T01:18:41Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [June 13, 2021, 1:18am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820/1 "2021-06-13T01:18:41Z")

</div>

Is there something like Python’s [`itertools.groupby`](https://docs.python.org/3/library/itertools.html#itertools.groupby), which is like Julia’s `Transducers.PartitionBy` but yields (partition\_key, partition\_entries) pairs instead of only partition entries?

cc @tkf

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 13, 2021, 2:10am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820/2 "2021-06-13T02:10:55Z")

</div>

Hmm… good point. I don’t remember why I didn’t pass along the key to the downstream transdcuers.

It is actually possible to write this, though:

```julia
julia> using Transducers, MicroCollections

julia> [1, 3, 2, 4, 3, 5] |>
       Map(x -> (isodd(x), x)) |>
       ReducePartitionBy(
           first,
           TeeRF(Map(first)'(right), Map(SingletonVector ∘ last)'(Completing(append!!))),
       ) |>
       collect
3-element Vector{Tuple{Bool, Vector{Int64}}}:
 (1, [1, 3])
 (0, [2, 4])
 (1, [3, 5])

```

(which is, BTW, parallelizable while PartitionBy is not)

OK, but arguably this is rather hairy to write.

Maybe it’d be better to wrap it in something like

```julia
reduced_partition_and_key(f, rf = Map(SingletonVector)'(Completing(append!!))) =
    Map(x -> (f(x), x)) |>
    ReducePartitionBy(
        first,
        TeeRF(Map(first)'(right), Map(last)'(rf)),
    )

```

so that

```julia
julia> [1, 3, 2, 4, 3, 5] |> reduced_partition_and_key(isodd) |> collect
3-element Vector{Tuple{Bool, Vector{Int64}}}:
 (1, [1, 3])
 (0, [2, 4])
 (1, [3, 5])

julia> [1, 3, 2, 4, 3, 5] |> reduced_partition_and_key(isodd, +) |> collect
3-element Vector{Tuple{Bool, Int64}}:
 (1, 4)
 (0, 6)
 (1, 8)

```

(The second example fuses in-partition reduction and avoids allocation of the inner vectors.)

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [June 13, 2021, 3:22am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820/3 "2021-06-13T03:22:05Z")

</div>

I don’t see why `Unique()` fails here. Replacing `Unique() |> collect` with `collect |> unique` works fine.

```julia
using Transducers, MicroCollections

reduced_partition_and_key(f, rf = Map(SingletonVector)'(Completing(append!!))) =
    Map(x -> (f(x), x)) |>
    ReducePartitionBy(
        first,
        TeeRF(Map(first)'(right), Map(last)'(rf)),
    )

charstrings = string.(collect("a123bc34d8ef34"))

charstrings |>
    reduced_partition_and_key(x->isnothing(tryparse(Int, x)), *) |>
    Filter(==(0) ∘ first) |>
    Map(x->parse(Int, x[2])) |>
    Unique() |>
    collect

ERROR: LoadError: MethodError: no method matching unwrap(::Transducers.Reduction{Unique{typeof(identity)},Transducers.Reduction{Map{Type{BangBang.NoBang.SingletonVector}},Transducers.BottomRF{Transducers.AdHocRF{typeof(BangBang.collector),typeof(identity),typeof(append!!),typeof(identity),Nothing}}}}, ::Tuple{Bool,String})

```

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 13, 2021, 4:52am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820/4 "2021-06-13T04:52:14Z")

</div>

Unfortunately, stateful transdcuers like `Unique` cannot be used after parallelizable transdcuer like `ReducePartitonBy`. It’s kind of a cost of parallelizability. There can be a better design to allow this but it’s a bit tricky to do ATM.

(Though the `unwrap` method error is actually a bug. Thanks for sharing the code!)

Meanwhile, I think the easiest approach might be to just cook up your `partitionby` using FGenerators:

```julia
julia> using FGenerators

julia> @fgenerator function partitionby(f, xs)
           buffer = eltype(xs)[]
           key = f(first(xs))
           for x in xs
               y = f(x)
               if !isequal(y, key)
                   @yield key => buffer
                   empty!(buffer)
                   key = y
               end
               push!(buffer, x)
           end
       end
partitionby (generic function with 1 method)

julia> partitionby(x->isnothing(tryparse(Int, x)), charstrings) |>
           Map(((k, v),) -> (k, prod(v))) |>
           Filter(==(0) ∘ first) |>
           Map(x->parse(Int, x[2])) |>
           Unique() |>
           collect
3-element Vector{Int64}:
 123
  34
   8

```

Note: `xs -> partitionby(f, xs)` is not a transducer so pre-processing of `xs` cannot be done with transdcuer.

---

<div class="post-metadata">

**Author:** ![tkf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkf/32/17635_2.png) [@tkf](https://discourse.julialang.org/u/tkf)\
**Post date:** [June 16, 2021, 6:28am UTC](https://discourse.julialang.org/t/partitionby-retaining-key/62820/5 "2021-06-16T06:28:17Z")

</div>

It just occurred to me that you’d need `isempty(buffer) || @yield key => buffer` at the end.
