# New syntax for slicing

**URL:** <https://discourse.julialang.org/t/new-syntax-for-slicing/91241>\
**Category:** Internals & Design\
**Tags:** proposal, arrays\
**Created:** [December 5, 2022, 2:54am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241 "2022-12-05T02:54:28Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 5, 2022, 2:54am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/1 "2022-12-05T02:54:28Z")

</div>

I’d like to suggest new syntax for `eachslice`, by overloading `getindex` (similar to how `:` behaves). Using `/` (which looks like a slice) would indicate a dimension to iterate over, i.e.

```julia
array[/, :] == eachslice(array; dims=1)

```

In addition to convenience and clarity, this should help users write more type-stable code. Some packages (like DimensionalData and AxisArrays) allow arbitrary axes for indexing into an array. Typically, these axes are included in the type. However, this means `eachslice` can be type-unstable, since the types of the slices depend on the value of `dims`. Constant propagation fixes this in most, but not quite all, cases (as we found when testing DimensionalData). Users slicing with `/` instead would get 100% guaranteed type stability, since the dimensions to be sliced across can be determined only from the types of the arguments.

Would the core devs be interested in an implementation of this?

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [December 5, 2022, 2:58am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/2 "2022-12-05T02:58:11Z")

</div>

This looks interesting. Could you put together a few examples showing this vs einstein notation for different common formulas?

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [December 5, 2022, 3:16am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/3 "2022-12-05T03:16:20Z")

</div>

You might like [StarSlice.jl](https://github.com/mcabbott/StarSlice.jl), which should probably be updated to the new eachslice.

(The idea of `*,:` came I believe from [early versions of](https://github.com/bramtayl/JuliennedArrays.jl/blob/v0.0.2/src/JuliennedArrays.jl) JuliennedArrays.)

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 5, 2022, 4:08pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/4 "2022-12-05T16:08:08Z")

</div>

As an example, squaring a matrix could be written:

```julia
map(sum, x[:, /], x[/, :]) == @einsum y[i, k] := x[i, j] * x[j, k]

```

I prefer the einsum notation in this case, but it isn’t as general. `x[:, /]` notation should have its advantage when working with operations that would be unnatural to write in Einstein notation. As an example, calculating the within-group variance for a panel dataset could be:

```julia
var.(x[time=/])

```

Rather than:

```julia
var.(eachslice(x; dims=:time))

```

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [December 5, 2022, 6:36pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/5 "2022-12-05T18:36:52Z")

</div>

Thats pretty cool syntax.

@ParadaCarleton the solution for DimensionalData.jl type stability is to use:

```julia
eachslice(array; dims=X())

```

As with most other base methods that accept `dims` keyword.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [December 5, 2022, 6:59pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/6 "2022-12-05T18:59:46Z")

</div>

This doesn’t necessarily need new syntax - `/` is a function/an object already after all, so it’d “just” need an overload of `getindex`. Might be a bit brittle though, if someone changes what `/` means in their package and doesn’t rely on the one exported from `Base`.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [December 5, 2022, 7:22pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/7 "2022-12-05T19:22:11Z")

</div>

> [@ParadaCarleton](#):
>
> `map(sum, x[:, /], x[/, :]) == @einsum y[i, k] := x[i, j] * x[j, k]`

I don’t think this works, there’s no method of `sum` taking two arrays, and the final container is 1-dimensional. Perhaps you meant either

```julia
slice(x) = broadcast(dot, eachslice(x; dims=2, drop=false), eachslice(x; dims=1, drop=false))
einsum(x) = @einsum y[i, k] := x[i, j] * x[j, k]
base(x) = x * x

```

or

```julia
slice2(x) = map(dot, eachslice(x; dims=2), eachslice(x; dims=1))
einsum2(x) = @einsum y[i] := x[i, j] * x[j, i]
base2(x) = vec(sum(x .* x'; dims=1)) # == diag(x * x)

```

> [@Sukera](#):
>
> This doesn’t necessarily need new syntax

See the package implementation above.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [December 5, 2022, 7:37pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/8 "2022-12-05T19:37:35Z")

</div>

> [@ParadaCarleton](#):
>
> Typically, these axes are included in the type. However, this means `eachslice` can be type-unstable, since the types of the slices depend on the value of `dims`. Constant propagation fixes this in most, but not quite all, cases

Potential eachslice(X; dims=Val(:lat)) is simple and typestable.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 6, 2022, 4:39pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/9 "2022-12-06T16:39:32Z")

</div>

Simple enough, but quite a bit longer than `x[lat=/]`.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 6, 2022, 4:49pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/10 "2022-12-06T16:49:57Z")

</div>

> [@mcabbott](#):
>
> I don’t think this works, there’s no method of `sum` taking two arrays, and the final container is 1-dimensional. Perhaps you meant

Yep, sorry for the error. So a possible notation for matrix multiplication using this new method would be:

```julia
dot.(x[:, /], x[/, :])

```

The goal of this isn’t to replace Einsum, though; I think Einstein notation is cleaner for tensor operations where it applies naturally. The main use case for `/` is wherever `eachslice` or `mapslices` is being used now, like summary statistics over slices.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [December 6, 2022, 5:01pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/11 "2022-12-06T17:01:37Z")

</div>

> [@ParadaCarleton](#):
>
> `var.(eachslice(x; dims=:time))`

Maybe worth noting that `var` has a method which takes `dims`. And this is much faster than working on individual rows, I think because it can pick a cache-friendly order in which to access the elements:

```julia
julia> let x = rand(100, 100)
         a = @btime var.(eachslice($x; dims=1)) # i.e. eachrow
         b = @btime var($x; dims=2)
         a ≈ vec(b)
       end
  16.250 μs (1 allocation: 896 bytes)
  3.833 μs (16 allocations: 2.44 KiB)
true

```

If you change `dims` so that this is `eachcol`, then the two are comparably fast. But if the goal is to handle arrays by dimension name, where you want not to care which dimension is first, then it seems you ought to worry about both (or all) cases.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 6, 2022, 8:55pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/12 "2022-12-06T20:55:40Z")

</div>

Using `var` with `dims` is a good solution for performance at the moment, but as mentioned elsewhere, having to implement `dims` keywords for all of these reducing functions by hand is both a bit limited and ugly design. It should be possible to implement the optimizations we currently do case-by-case for `sum`, `var`, etc. in a more general way, but it’s not immediately clear to me how.

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 7, 2022, 2:00am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/13 "2022-12-07T02:00:18Z")

</div>

I think the big optimization here is these are all reducing functions, so you can write the same operation as either `map(f, eachslice(x; dims))` or as `reduce(f, x; dims=Not(dims))`\*. Sometimes, memory-locality means that rewriting this way can be a huge improvement.

A while back, I thought it would be really nice if traits were added to the language, and could be added to functions as well as structs. If that were the case, marking reducing functions like `sum` and `var` with a `Reduces` trait would let us take advantage of this reorder-for-memory-locality optimization by dispatching on the function trait.

\*this doesn’t work with `var` itself, but variance can be rewritten as a reducing function, which is what I assume the implementation with the `dims` keyword is doing internally.

---

<div class="post-metadata">

**Author:** ![mcabbott](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mcabbott/32/6603_2.png) [@mcabbott](https://discourse.julialang.org/u/mcabbott)\
**Post date:** [December 7, 2022, 3:33am UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/14 "2022-12-07T03:33:41Z")

</div>

> [@ParadaCarleton](#):
>
> what I assume the implementation with the `dims` keyword is doing

It leads to [`centralize_sumabs2!`](https://github.com/JuliaStats/Statistics.jl/blob/e9ac70b760dcf87b77affe6c068548a3325d6e2b/src/Statistics.jl#L248-L290), which looks a lot like a re-implementation of reduce. I guess it’s not literally `reduce` as it has to index the array of pre-computed means at the same time as the input & output arrays. Are you suggesting a tidier way to implement that?

The case without dims is literally `mapreduce(centralizedabs2fun(m), +, A)` (just above). And this is what’s called by any map over slices, of course. It pre-computes the mean, whereas [OnlineStats.jl](https://github.com/joshday/OnlineStats.jl/blob/1fca9957372d6120bf80f11bca56356d3a7e3609/src/stats/kahan.jl#L112-L129) does it in one pass, but seems slower.

I’m not sure if we’re talking about a notation for slicing anymore, though, so much as how best to (a) implement various reduction-like things efficiently, and (b) how best to call them.

For (a) it seems clear that the implementation cannot always just work on slices. For (b), one objection to an API which pretends you are working with slices, while actually doing something different, is that this seems a little fragile. It’s a magic fast path a little like the `reduce(vcat, xs)` one – which has [surprising](https://discourse.julialang.org/t/how-can-i-do-list-comprehension-for-vcat/84178/11) ways to step off the fast path accidentally. With the “ugly design” of having a `dims` method _if and only if_ there exists a faster-than-slices implementation, at least you know what you’re getting.

I assume that’s the implicit rule in Base/std.lib. But might be wrong; are there examples of things with `dims` methods which do just always call `mapslices` or something? (Is seems that [`median`](https://github.com/JuliaStats/Statistics.jl/blob/e9ac70b760dcf87b77affe6c068548a3325d6e2b/src/Statistics.jl#L866) does call mapslices, but uses mutation to do a bit better than `mapslices(median, x; dims)`.)

> [@ParadaCarleton](#):
>
> implement the optimizations we currently do case-by-case for `sum`, `var`, etc. in a more general way

Quite a few do share the same implementation as `sum`. Maybe it would help to list others you include in “etc”?

---

<div class="post-metadata">

**Author:** ![ParadaCarleton](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/paradacarleton/32/20005_2.png) [@ParadaCarleton](https://discourse.julialang.org/u/ParadaCarleton)\
**Post date:** [December 7, 2022, 6:30pm UTC](https://discourse.julialang.org/t/new-syntax-for-slicing/91241/15 "2022-12-07T18:30:26Z")

</div>

> [@mcabbott](#):
>
> Quite a few do share the same implementation as `sum`. Maybe it would help to list others you include in “etc”?

`sum`, `prod`, `all`, `any`, `⊻`, `⊼`, `maximum`, `minimum`, `median`, `mode`, and every other kind of summary statistic.

> [@mcabbott](#):
>
> Are you suggesting a tidier way to implement that?

You could definitely implement it to calculate both means and variances in one pass, and then reduce using mean(var)+var(mean), but that’s getting off-topic for sure.

> [@mcabbott](#):
>
> For (a) it seems clear that the implementation cannot always just work on slices. For (b), one objection to an API which pretends you are working with slices, while actually doing something different, is that this seems a little fragile. It’s a magic fast path a little like the `reduce(vcat, xs)` one – which has [surprising](https://discourse.julialang.org/t/how-can-i-do-list-comprehension-for-vcat/84178/11) ways to step off the fast path accidentally. With the “ugly design” of having a `dims` method _if and only if_ there exists a faster-than-slices implementation, at least you know what you’re getting.

That’s exactly what I’m proposing we should fix with the proposal I made. The current `dims` keyword argument is exactly that kind of magic fastpath–using a `dims` keyword argument makes this dramatically faster than the fully-supported `map(sum, eachslice(x; dims))` syntax and stands in contrast to the always-fast `reduce(+, x; dims))` (which is just another syntax for `sum`). (To be clear, I don’t think `dims` is always ugly; I just think it’s ugly having to reimplement it for half-a-dozen different functions case-by-case.) The `reduce` syntax is better and should be encouraged, but there’s nothing wrong with trying to optimize `map` so new users don’t find their code slow if they try using `eachslice` instead.
