# Find filtered min/max index

**URL:** https://discourse.julialang.org/t/find-filtered-min-max-index/95154
**Category:** General Usage
**Tags:** question
**Created:** [February 24, 2023, 5:16pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154 "2023-02-24T17:16:25Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![filchristou](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/filchristou/32/26760_2.png) [@filchristou](https://discourse.julialang.org/u/filchristou)
#### Post date: [February 24, 2023, 5:16pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/1 "2023-02-24T17:16:25Z")

</div>

Every now and then I need to find the index of the min/max element of a filtered data structure.  
I was wondering, how do you normally tackle this quite common problem?

As a first reaction I think of `findmin(o, filter(f, data))` but filtering doesn’t hold the original indices.  
I ended up using `findmin(o, skipmissing([f(d) ? d : misssing for d in data])` where I inject `missing `values in the values that should be filtered out.

MWE:

```julia
julia> using Random

julia> rng = MersenneTwister(0);

julia> data = [(rand(rng), rand(rng)) for _ in 1:10];

julia> dataobjective(d) = d[2];

julia> filterdata(d) = d[1] > 0.5;

julia> findmin(dataobjective, skipmissing([filterdata(d) ? d : missing for d in data]))
(0.5392892841426182, 6)

```

It’s very convenient that `skipmissing` is propagating the original indices. Is there something like `skipfilter` that would do exactly as `skipmissing` but for elements satisfying a predicate ?  
And how would you solve the above problem ?

---

<div class="post-metadata">

### Author: ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)
#### Post date: [February 24, 2023, 5:53pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/2 "2023-02-24T17:53:52Z")

</div>

Check out @aplavin’s Skipper.jl, which provides an arbitrary-predicate `skip` that should totally be in Base imo.

---

<div class="post-metadata">

### Author: ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)
#### Post date: [February 24, 2023, 6:11pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/3 "2023-02-24T18:11:17Z")

</div>

```julia
minimum(t for t in zip(last.(data),eachindex(data)) if t[1]>0.5)

```

the elegant @lmiq solution applied to the (modified 😄 )OP case

```julia
findmin(x -> x[2] > 0.5 ? x[2] : +Inf, data) 

```

---

<div class="post-metadata">

### Author: ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)
#### Post date: [February 24, 2023, 6:11pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/4 "2023-02-24T18:11:38Z")

</div>

You can use the predicate to filter, for example:

```julia
julia> data = collect(1:10)
10-element Vector{Int64}:
  1
  2
  3
  4
  5
  6
  7
  8
  9
 10

julia> findmin(x -> x > 3 ? x : +Inf, data)
(4, 4)

```

or, similar to the OP example:

```julia
julia> data = [(rand(),rand()) for _ in 1:10];

julia> findmin(d -> d[1] > 0.5 ? d[2] : +Inf, data)
(0.42413123524471796, 8)

```

---

<div class="post-metadata">

### Author: ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)
#### Post date: [February 24, 2023, 6:40pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/5 "2023-02-24T18:40:10Z")

</div>

okay. I read the OP more carefully

```julia
minimum((t[1][2],t[2]) for t in zip(data, eachindex(data)) if t[1][1] > 0.5)

```

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [February 24, 2023, 7:14pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/6 "2023-02-24T19:14:58Z")

</div>

Link?

---

<div class="post-metadata">

### Author: ![filchristou](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/filchristou/32/26760_2.png) [@filchristou](https://discourse.julialang.org/u/filchristou)
#### Post date: [February 24, 2023, 7:22pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/7 "2023-02-24T19:22:07Z")

</div>

> **[Alexander Plavin / Skipper.jl · GitLab](https://gitlab.com/aplavin/Skipper.jl)**
>
> GitLab.com

I also had problem finding it with google/bing. It appears they don’t parse gitlab that well as github.

---

<div class="post-metadata">

### Author: ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)
#### Post date: [February 24, 2023, 7:24pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/8 "2023-02-24T19:24:38Z")

</div>

How about:

```julia
julia> withmissing = [filterdata(d) ? d : missing for d in data]

julia> findmin(map(withmissing) do x 
    ismissing(x) ? (true,0.0) : (false, dataobjective(x)) 
  end)
((false, 0.5392892841426182), 6)

```

This generalizes well to other cases, but needs a tiny addition to access the value. More importantly, it handles the all missing case which could potentially cause trouble.

Or, same as rocco\_sprmnt21 's but perhaps more readable:

```julia
minimum( [ (dataobjective(x),i) for (i,x) in enumerate(withmissing)
          if !ismissing(x) ] )

```

---

<div class="post-metadata">

### Author: ![lmiq](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lmiq/32/18314_2.png) [@lmiq](https://discourse.julialang.org/u/lmiq)
#### Post date: [February 24, 2023, 7:44pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/9 "2023-02-24T19:44:43Z")

</div>

That should definitely go to base. Why it is not?

Just checked it out:

```julia
julia> using Skipper

julia> data = collect(1:10);

julia> findmin(skip(x -> x < 4, data))
(4, 4)

julia> @btime findmin(skip(x -> x < 4, $data))
  21.599 ns (0 allocations: 0 bytes)
(4, 4)

```

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [February 24, 2023, 7:49pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/10 "2023-02-24T19:49:10Z")

</div>

Aside from inverting the predicate, what is the difference between `skip` and `Iterators.filter`?

---

<div class="post-metadata">

### Author: ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)
#### Post date: [February 24, 2023, 8:06pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/11 "2023-02-24T20:06:34Z")

</div>

The intent of `filter` is to construct an entirely new collection excluding the filtered elements, and `Iterators.filter` is a lazy variant.

The intent of `skipmissing` (and `skip`) is different: we intend to continue working with the same collection, but we’re specifying that for this operation certain elements won’t be under consideration.

The example here is a good one for demonstrating the difference; because `filter` returns a brand new collection, the indices will be entirely different from the original collection. In reflection of this, `Iterators.filter` isn’t even indexable. `skip`, on the other hand, does what we want here.

```julia
julia> using Skipper
       data = collect(1:10)
       findmin(skip(<(4), data))
(4, 4)

julia> findmin(Iterators.filter(!<(4), data))
ERROR: MethodError: no method matching keys(::Base.Iterators.Filter{ComposedFunction{typeof(!), Base.Fix2{typeof(<), Int64}}, Vector{Int64}})

```

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [February 24, 2023, 8:16pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/12 "2023-02-24T20:16:09Z")

</div>

In other words, `skip` makes a view while `filter` and `Iterators.filter` makes a copy?

---

<div class="post-metadata">

### Author: ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)
#### Post date: [February 24, 2023, 9:11pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/13 "2023-02-24T21:11:03Z")

</div>

Sort of. I’d say `skip` creates a view of the original collection. In contrast, `filter` creates an entirely new (copied and modified) collection, and `Iterators.filter` creates a view of what the new collection will look like without actually collecting it.

When I think about it though, outside from the fact that we want generators to resemble comprehensions, I can’t think of any cases that you’d use `Iterators.filter` where you couldn’t use `skip`, but I can think of plenty of cases where you’d use `skip` where you couldn’t use `Iterators.filter`. Maybe a lazy `filter` was never quite the right idea to begin with?

Either that, or maybe the wrong idea is actually eager `filter`: by implicitly creating a new collection it has trained us to believe a filtered collection is a new collection—maybe `Base.filter` should actually be lazy and act like `skip` with its predicate inversed, and you’d have to explicitly call `collect` for the indices to change. That would be more efficient when building filter-\>map-\>reduce chains, and by preserving indices it’d also make `skip` unnecessary.

Definitely brings to mind @jakobnissen’s [complaints about Julia’s FP primitives](https://viralinstruction.com/posts/badjulia/#functional_programming_primitives_are_not_well_designed).

> [@lmiq](#):
>
> Why it is not?

I don’t think anybody has opened an issue or PR.

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [February 24, 2023, 9:18pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/14 "2023-02-24T21:18:21Z")

</div>

I’ll just mention the opposite of `skip` could be `keep`. I know some people find the name “filter” confusing since it sounds like it’s filtering _out_ but it’s the opposite.

---

<div class="post-metadata">

### Author: ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)
#### Post date: [February 24, 2023, 10:22pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/15 "2023-02-24T22:22:46Z")

</div>

Would it be possible to transition Julia’s generator and comprehension syntax to lower to a `Keep` type instead of a `Filter` type?

If anything could justify a Julia 2.0, it might be the `Base` FP primitives. They’re FUBAR and guide people into all sorts of performance traps. Hard to claim your language is performant, if the base primitives guide the ecosystem into making slow code. And for a language whose _raison d’être_ is performance, that’s a matter of existential import.

As an aside, `collect` is FUBAR too. Its current role should be satisfied by the `Array` constructor, and then `collect` could be used for collecting all sorts of other collection types.

---

<div class="post-metadata">

### Author: ![jar1](https://avatars.discourse-cdn.com/v4/letter/j/c0e974/32.png) [@jar1](https://discourse.julialang.org/u/jar1)
#### Post date: [February 24, 2023, 10:32pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/16 "2023-02-24T22:32:06Z")

</div>

If that happens, I’d also consider changing the Base functions to enable [error handling](https://github.com/tkf/Try.jl) and possibly passing memory allocators.

---

<div class="post-metadata">

### Author: ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)
#### Post date: [February 25, 2023, 11:06am UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/17 "2023-02-25T11:06:15Z")

</div>

> [@filchristou](#):
>
> I need to find the index of the min/max element of a filtered data structure

To find only the index, there is `argmin`.  
Using the functions in the OP:

```julia
# 22 ns (0 allocs: 0 bytes)
argmin(filterdata(d) ? dataobjective(d) : +Inf for (_,d) in pairs(data))

```

---

<div class="post-metadata">

### Author: ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)
#### Post date: [February 25, 2023, 4:45pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/18 "2023-02-25T16:45:19Z")

</div>

> [@jar1](#):
>
> In other words, `skip` makes a view while `filter` and `Iterators.filter` makes a copy?

`Iterators.filter` is also lazy. The difference is that `filter`s don’t keep the original indices, while `skipmissing` in `Base` and `skip` in `Skipper` do keep. `filter` cannot keep original indices by design.

---

<div class="post-metadata">

### Author: ![rocco\_sprmnt21](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rocco_sprmnt21/32/20127_2.png) [@rocco\_sprmnt21](https://discourse.julialang.org/u/rocco_sprmnt21)
#### Post date: [February 25, 2023, 5:03pm UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/19 "2023-02-25T17:03:08Z")

</div>

regarding performance the following version

```julia
last(reduce((s,c)->(c[1]>.5 && c[2]<s[2]) ? (s[1]+1,c[2],s[1]+1) : (s[1]+1, s[2],s[3]), data, init=(0,Inf,0))) 

```

seems to rival even a for loop form

```julia

julia> @btime last(reduce((s,c)->(c[1]>.5 && c[2]<s[2]) ? (s[1]+1,c[2],s[1]+1) : (s[1]+1, s[2],s[3]), $data, init=(0,Inf,0)))
  10.000 ns (0 allocations: 0 bytes)
3

julia> @btime argmin(i -> filterdata($data[i]) ? dataobjective($data[i]) : +Inf, axes($data,1))
  23.370 ns (0 allocations: 0 bytes)
3

julia> data = [(rand(rng), rand(rng)) for _ in 1:10^6];    

julia> @btime argmin(i -> first($data[i])>0.5 ? last($data[i]) : +Inf, axes($data,1))
  4.612 ms (0 allocations: 0 bytes)
669781

julia> @btime last(reduce((s,c)->(c[1]>.5 && c[2]<s[2]) ? (s[1]+1,c[2],s[1]+1) : (s[1]+1, s[2],s[3]), $data, init=(0,Inf,0)))
  1.784 ms (0 allocations: 0 bytes)
669781

julia> @btime argmin(i -> filterdata($data[i]) ? dataobjective($data[i]) : +Inf, axes($data,1))
  4.668 ms (0 allocations: 0 bytes)
669781

```

```julia
function ffm(arr)
    i,im, min=0, 0, Inf
    for e in arr
        i+=1
        if e[1]>.5 && e[2]<min
            im, min = i, e[2]
        end
    end
    im
end
julia> @btime ffm($data)
  1.804 ms (0 allocations: 0 bytes)
669781

```

---

<div class="post-metadata">

### Author: ![uniment](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/uniment/32/24532_2.png) [@uniment](https://discourse.julialang.org/u/uniment)
#### Post date: [February 26, 2023, 11:12am UTC](https://discourse.julialang.org/t/find-filtered-min-max-index/95154/20 "2023-02-26T11:12:41Z")

</div>

Is there a running list of things that should be fixed in 2.0? And is there a Julia fork that would allow package devs to start porting their code over to it?

> [@aplavin](#):
>
> `filter` cannot keep original indices by design.

If you were designing this interface from scratch, starting with the lazy methods instead of the eager methods, would you make `filter` keep the original indices?

[Next page](https://discourse.julialang.org/t/find-filtered-min-max-index/95154.md?page=2)
