# Request to upgrade to LossFunctions.jl

**URL:** <https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738>\
**Category:** Machine Learning\
**Tags:** package, release, machine-learning\
**Created:** [April 21, 2023, 12:06pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738 "2023-04-21T12:06:29Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 21, 2023, 12:06pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/1 "2023-04-21T12:06:29Z")

</div>

Dear ML community, we released [LossFunctions.jl](https://github.com/JuliaML/LossFunctions.jl) v0.9 with a few important breaking changes:

- Reversed order of arguments to match other ecosystems, loss(yhat, y) is now the order.
- Removed the ObsDim business to support a more general interface with iterables of observations.
- Removed OrdinalMarginLoss to support a more general interface with CategoricalArrays.jl].
- Dropped unnecessary dependencies, the only dependency is CategoricalArrays.jl now.

These changes will allows us to sync the widely tested loss functions in LossFunctions.jl with loss functions defined in other ecosystems such as Flux.jl. I would like to invite every ML contributor in Julia to try to move their own loss function implementations to LossFunctions.jl and join efforts maintaining the package.

Our next goal is to formalize the support for general datasets with known observation dimension (e.g. Tables.jl). Previously datasets were assumed to be n-dimensional arrays, which is quite limiting and low-level.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 21, 2023, 5:54pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/2 "2023-04-21T17:54:44Z")

</div>

Wonder what’s the motivation for implicit broadcasting? Julia doesn’t really do that elsewhere…

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 21, 2023, 6:00pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/3 "2023-04-21T18:00:58Z")

</div>

Can you elaborate?

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 21, 2023, 6:04pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/4 "2023-04-21T18:04:35Z")

</div>

From the docs:

```julia
julia> value(L2DistLoss(), 1.0, 0.5)
0.25

julia> value(L2DistLoss(), true_targets, pred_outputs)
3-element Array{Float64,1}:
 0.25
 4.0
 1.0

```

Following Julian approach, the latter should be `value.(...)`, with a dot.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 21, 2023, 6:51pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/5 "2023-04-21T18:51:41Z")

</div>

In this case we want to preserve the vectorized version in order to optimize aggregation methods. Check the AggMode examples where the result is usually summed up.

Optimizations can dispatch on different types of dataset to aggregate the scalar version more efficiently.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 22, 2023, 6:52pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/6 "2023-04-22T18:52:42Z")

</div>

Aggregation is a different story, I was only talking about the `value` method that takes and returns arrays. Julia specifically avoids defining “implicitly broadcasted” methods, as there is no reason to do this.

Aggregation is completely orthogonal to that. Although, instead of

```julia
value(L1DistLoss(), [1,2,3], [2,5,-2], AggMode.Sum())

```

one of the following would be cleaner and not require learning new symbols/objects (`AggMode`, `Sum`):

```julia
# minimal change from yours:
value(L1DistLoss(), [1,2,3], [2,5,-2], sum)
# like Julia Base map(f, [...], [...]), but aggregate:
sum(L1DistLoss(), [1,2,3], [2,5,-2])

```

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 23, 2023, 12:15pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/7 "2023-04-23T12:15:21Z")

</div>

I like the proposed alternatives and we could consider them in the next breaking release. The only thing I don’t see yet is how to incorporate weights in these summations? It is easy when the losses are summed up without weights as we can simply rely on `sum` and `mean` from Base.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 23, 2023, 12:30pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/8 "2023-04-23T12:30:31Z")

</div>

> [@juliohm](#):
>
> how to incorporate weights in these summations?

You can follow `StatsBase` and do

```julia
mean(L1DistLoss(), [1,2,3], [2,5,-2], weights([1, 2, 1]))

```

Also: see the already existing `dims` argument of `sum`/`mean` instead of introducing `ObsDim`.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 23, 2023, 12:39pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/9 "2023-04-23T12:39:10Z")

</div>

I think you missed the point of this release. We already removed ObsDim, and are continuously updating the code to look more Julian:

> <https://github.com/JuliaML/LossFunctions.jl/issues/126>
>
> Dear all,
> 
> I would like to propose a major refactoring of the codebase to \*sim…plify\* future additions and generalizations, and to facilitate future contributions. I think we are overusing macros in the package with the only purpose of sharing docstrings among various losses, and we could eliminate this entry barrier to potential contributors.
> 
> In particular, I would like to suggest a few modifications here, and ask for your approval before I start submitting PRs.
> 
> \# Suggestions of improvement
> 
> 1. Can we get rid of the \`value\_fun\`, \`deriv\_fun\`, \`deriv2\_fun\`, \`value\_deriv\_fun\` functionality? I understand that these functions were created in the past because the language didn't have efficient lambdas and closures, etc. Moving forward, I think we could stick to a single interface for evaluating losses \`value\`, \`deriv\`, and \`deriv2\` where the last two functions could have fallback implementations via auto-diff when the user only implements the \`value\` function.
> 
> 2. Similarly, can we get rid of the following functor syntax in the top source file, and stick to the single API defined above?
> 
> \`\`\`julia
> \# allow using some special losses as function
> (loss::ScaledSupervisedLoss)(args...) = value(loss, args...)
> (loss::WeightedBinaryLoss)(args...) = value(loss, args...)
> 
> \# allow using SupervisedLoss as function
> for T in filter(isconcretetype, subtypes(SupervisedLoss))
> @eval (loss::$T)(args...) = value(loss, args...)
> end
> 
> \# allow using MarginLoss and DistanceLoss as function
> for T in union(subtypes(DistanceLoss), subtypes(MarginLoss))
> @eval (loss::$T)(args...) = value(loss, args...)
> end
> \`\`\`
> 
> 3. Can we get rid of the internal (no-exported) types \`Deriv\` and \`Deriv2\`? I understand that they are only used internally for plotting derivatives in \`src/supervised/io.jl\`.
> 
> 4. Can we simplify the test suite? Currently, it seems to be testing the same functionality with hundreds of numbers giving the illusion of good coverage and making any run of the tests take forever when tests fail (IO bottlenecks).
> 
> 5. Can we simplify the loop in \`src/supervised/supervised.jl\` that loops over \`value\`, \`deriv\` and \`deriv2\`? In particular, I am curious if we could only define (without metaprogramming) the aggregation of \`value\`, and then rely on auto-diff to compute \`deriv\` and \`deriv2\`? This is a modification that we need to think more carefully, but that could simplify the codebase tremendously. If auto-diff does not work for all losses, we can always provide a specific implementation to overwrite the auto-diff defaults. My question is, can auto-diff be performed at compile time? Do we get any slowdown if we follow this design?
> 
> I will start working on separate PRs for items (1) (2) (3) and (5). I need your input before I can start working on (5).

We will consider the StatsBase.jl approach for weights in our next brainstorming phase.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 23, 2023, 12:58pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/10 "2023-04-23T12:58:31Z")

</div>

> [@juliohm](#):
>
> I think you missed the point of this release. We already removed ObsDim

Nice, so that must be a docs issue? I’ve been looking at the “latest” docpage at [Efficient Sum and Mean · LossFunctions.jl](https://juliaml.github.io/LossFunctions.jl/latest/user/aggregate/), it should be there right? As the release already happened:

> [@juliohm](#):
>
> we released [LossFunctions.jl](https://github.com/JuliaML/LossFunctions.jl) v0.9

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 23, 2023, 1:07pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/11 "2023-04-23T13:07:55Z")

</div>

Yes, probably a docs issue. We updated the docs but the build scripts did not deploy it apparently 😕

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [April 23, 2023, 4:16pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/12 "2023-04-23T16:16:12Z")

</div>

> [@aplavin](#):
>
> Wonder what’s the motivation for implicit broadcasting?

Most machine learning metrics do both a mapping and a reduction in order to calculate a scalar metric. I think it’s reasonable to encapsulate that in one function with a signature like `mymetric(ŷ, y)`, where `ŷ` and `y` are iterables. And in fact that is what both MLJ and Flux do:

[https://alan-turing-institute.github.io/MLJ.jl/dev/performance\_measures/#Using-built-in-measures](https://alan-turing-institute.github.io/MLJ.jl/dev/performance_measures/#Using-built-in-measures)

[https://fluxml.ai/Flux.jl/stable/models/losses/](https://fluxml.ai/Flux.jl/stable/models/losses/)

Personally, I don’t really see a need for a `value` function. It seems like regular old explicit functions like `mae(ŷ, y)` are good enough and a lot easier to read.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 12:13pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/13 "2023-04-24T12:13:30Z")

</div>

We are refactoring the package as a whole, it is likely that the next release will get rid of the AggMode submodule and will use a more Julian approach.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 8:20pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/14 "2023-04-24T20:20:17Z")

</div>

A big advantage of the AggMode currently implemented in LossFunctions.jl is that it doesn’t allocate intermediate arrays. We can dispatch on specific aggregation methods and reduce the terms of the aggregation without broadcasting a big vector of losses to be summed up later. Am I missing something?

Started cleaning up AggMode.None in this PR:

> <https://github.com/JuliaML/LossFunctions.jl/pull/162>
>
> Continuing the refactoring discussed in #126. This PR simplifies the machinery a…round aggregation of losses over a dataset. The goal is to rely on simple Julia reductions as much as possible.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 8:32pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/15 "2023-04-24T20:32:16Z")

</div>

> [@CameronBieganek](#):
>
> [Performance Measures · MLJ](https://alan-turing-institute.github.io/MLJ.jl/dev/performance_measures/#Using-built-in-measures)
> 
> [https://fluxml.ai/Flux.jl/stable/models/losses/](https://fluxml.ai/Flux.jl/stable/models/losses/)

Thanks for pointing these out @CameronBieganek. The main difference I see between MLJ and Flux is that MLJ decided to place weight vectors as the third argument of the `loss(ŷ, y, w)`, whereas Flux decided to place the weights in the aggregation function `loss(ŷ, y, agg=x->mean(w .* x))`. Moreover, MLJ defaults to no aggregation whereas Flux defauls to `mean`.

In terms of performance, the MLJ approach gives more opportunity to avoid memory allocations. For example, if the aggregation is `sum` we can simply

`sum(wi * loss(ŷi, yi) for (ŷi, yi, wi) in zip(ŷ, y, w))`

without ever allocating intermediate arrays.

I am tempted to implement the MLJ approach, but am open to mroe input before proceeding.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 8:36pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/16 "2023-04-24T20:36:11Z")

</div>

I am starting to consider that maybe the AggMode.Sum, AggMode.Mean and AggMode.WeightedSum should be preserved as types in order to dispatch more efficient implementations for the different types of aggregations. Would be happy to be convinced otherwise.

---

<div class="post-metadata">

**Author:** ![CameronBieganek](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/cameronbieganek/32/6915_2.png) [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Post date:** [April 24, 2023, 10:06pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/17 "2023-04-24T22:06:03Z")

</div>

> [@juliohm](#):
>
> Moreover, MLJ defaults to no aggregation

Well, many MLJ metrics use aggregation, but some do not. They have a trait `reports_each_observation` that specifies whether or not aggregation is used.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 10:12pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/18 "2023-04-24T22:12:44Z")

</div>

Thanks for clarifying. We will certainly not follow this pattern. We want the same behavior always and extra options to change the behavior to something else if necessary.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [April 24, 2023, 11:18pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/19 "2023-04-24T23:18:49Z")

</div>

I think the “implicit broadcasting” methods (like `L1DistLoss()([1,2,3], [2,5,-2])`) should just be removed, and users directed to regular Julia broadcasting.  
This change doesn’t affect aggregations at all!

Regarding aggregations:

> [@juliohm](#):
>
> A big advantage of the AggMode currently implemented in LossFunctions.jl is that it doesn’t allocate intermediate arrays.

Big advantage with respect to what?  
Surely you can dispatch stuff like

```julia
sum(L1DistLoss(), [1,2,3], [2,5,-2])
mean(L1DistLoss(), [1,2,3], [2,5,-2], weights([1, 2, 1]))

```

to exactly the same kind of code as the current `value(L1DistLoss(), [1,2,3], [2,5,-2], AggMode.Sum())` does. With exact same performance.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [April 24, 2023, 11:24pm UTC](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738/20 "2023-04-24T23:24:09Z")

</div>

> [@aplavin](#):
>
> Surely you can dispatch stuff like
> 
> ```julia-auto
> sum(L1DistLoss(), [1,2,3], [2,5,-2])
> mean(L1DistLoss(), [1,2,3], [2,5,-2], weights([1, 2, 1]))
> 
> ```
> 
> to exactly the same kind of code as the current `value(L1DistLoss(), [1,2,3], [2,5,-2], AggMode.Sum())` does. With exact same performance.

I think we can assume that these are the only two useful aggregation functions, i.e. `sum` and `mean` with optional weights, and then get rid of all aggregation types in LossFunctions.jl. Thanks for the suggestion.

[Next page](https://discourse.julialang.org/t/request-to-upgrade-to-lossfunctions-jl/97738.md?page=2)
