# Softmax and large numbers

**URL:** <https://discourse.julialang.org/t/softmax-and-large-numbers/58674>\
**Category:** General Usage\
**Tags:** flux\
**Created:** [April 6, 2021, 10:50am UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674 "2021-04-06T10:50:15Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![gdkrmr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdkrmr/32/2791_2.png) [@gdkrmr](https://discourse.julialang.org/u/gdkrmr)\
**Post date:** [April 6, 2021, 10:50am UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/1 "2021-04-06T10:50:15Z")

</div>

I need to apply a softmax layer to values that are potentially large and I am getting `NaN` values. Playing around a bit more I found the following:

```julia
julia> using CUDA, Flux

julia> softmax(gpu(Float32[1, 2, 3, Inf]))
4-element CuArray{Float32,1}:             
   0.0                                    
   0.0                                    
   0.0                                    
 NaN                                      
                                          
julia> softmax(Float32[1, 2, 3, Inf])     
4-element Array{Float32,1}:               
 0.0                                      
 0.0                                      
 0.0                                      
 1.0                                      

```

For comparison, both tensorflow and pytorch return `[nan, nan, nan, nan]` for this on cpu and gpu. Is this a bug? What is the “correct” implementation?

---

<div class="post-metadata">

**Author:** ![DNF](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dnf/32/10191_2.png) [@DNF](https://discourse.julialang.org/u/DNF)\
**Post date:** [April 6, 2021, 12:43pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/2 "2021-04-06T12:43:50Z")

</div>

It seems like the implementations for CUDA and CPU arrays are different. The cpu version explicitly handles infinities, while cuda does not:  
[https://github.com/FluxML/NNlib.jl/blob/master/lib/NNlibCUDA/src/cudnn/softmax.jl](https://github.com/FluxML/NNlib.jl/blob/master/lib/NNlibCUDA/src/cudnn/softmax.jl)  
vs

> <https://github.com/FluxML/NNlib.jl/blob/2c8af3051cb6b41f8b07f3357edaaed1b5db2f8e/src/softmax.jl#L57>

If you look at line 57 you see that Inf is handled specially.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 6, 2021, 1:12pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/3 "2021-04-06T13:12:06Z")

</div>

> [@gdkrmr](#):
>
> What is the “correct” implementation?

`[0,0,0,1]` is unambiguously the correct output, as you can easily see if you compute

\lim\_{x\to\infty} \frac{[e^1, e^2, e^3, e^x]}{e^1 + e^2 + e^3 + e^x}

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 6, 2021, 1:17pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/4 "2021-04-06T13:17:44Z")

</div>

> [@DNF](#):
>
> ```julia
> function softmax!(out::AbstractArray{T}, x::AbstractArray; dims = 1) where {T}
> max_ = maximum(x; dims = dims)
> if all(isfinite, max_)
> out .= exp.(x .- max_)
> else
> @. out = ifelse(isequal(max_,Inf), ifelse(isequal(x,Inf), 1, 0), exp(x - max_))
> end
> out ./= sum(out; dims = dims) # could re-use max_ when dims != (:) and eltype(x) == T.
> end
> 
> ```

This seems mathematically questionable to me in the case where there are multiple `Inf` entries, because it assumes that all `Inf` values are the same (i.e. it acts as though `Inf/Inf == 1`).

For example, it gives `softmax([1,2,Inf,Inf]) == [0,0,0.5,0.5]`, whereas I would tend to say that a more formally correct answer would be `[0,0,NaN,NaN]`.

That being said, perhaps it is more useful in ML applications to give `0.5` than `NaN`, i.e. to split the `softmax` result equally between all `Inf` entries.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [April 6, 2021, 1:40pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/5 "2021-04-06T13:40:16Z")

</div>

> [@gdkrmr](#):
>
> I need to apply a softmax layer to values that are potentially large and I am getting `NaN` values.

For `softmax` there’s a difference between “potentially large” and infinite. In the latter case, as discussed in other replies, you have to special case the implementation and in the case of multiple infinities resort to conventions.

For “potentially large”, on the other hand, you will run into trouble with a naive implementation. E.g.

```julia
julia> x=Float32.([87, 88, 89, 90])
4-element Vector{Float32}:
 87.0
 88.0
 89.0
 90.0

julia> exp.(x) ./ sum(exp.(x))
4-element Vector{Float32}:
   0.0
   0.0
 NaN
 NaN

```

since

```julia
julia> exp.(x)
4-element Vector{Float32}:
  6.0760303f37
  1.6516363f38
 Inf
 Inf

```

The canonical solution to this is to first subtract the largest value from all elements, as you can also see in the pasted code in another reply.

```julia
julia> y = x .- maximum(x)
4-element Vector{Float32}:
 -3.0
 -2.0
 -1.0
  0.0

julia> exp.(y)
4-element Vector{Float32}:
 0.049787067
 0.13533528
 0.36787945
 1.0

julia> exp.(y) ./ sum(exp.(y))
4-element Vector{Float32}:
 0.032058604
 0.08714432
 0.23688284
 0.6439143

```

This way all the exponentiated values are scaled proportionally so that the largest value is one and the overflow problems are gone. You might _underflow_ the smaller elements but that has no practical consequence.

---

<div class="post-metadata">

**Author:** ![GunnarFarneback](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gunnarfarneback/32/1827_2.png) [@GunnarFarneback](https://discourse.julialang.org/u/GunnarFarneback)\
**Post date:** [April 6, 2021, 1:57pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/6 "2021-04-06T13:57:23Z")

</div>

> [@stevengj](#):
>
> That being said, perhaps it is more useful in ML applications to give `0.5` than `NaN` , i.e. to split the `softmax` result equally between all `Inf` entries.

Splitting it can be a reasonable graceful degradation and if it’s part of a model that is trained end-to-end it may plausibly learn to take the convention into account. However, most of the time something has gone wrong if you run into infinities at all, and a `NaN` output will make the problem more apparent.

---

<div class="post-metadata">

**Author:** ![gdkrmr](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdkrmr/32/2791_2.png) [@gdkrmr](https://discourse.julialang.org/u/gdkrmr)\
**Post date:** [April 6, 2021, 7:35pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/7 "2021-04-06T19:35:49Z")

</div>

Thanks everyone, this discussion was really insightful!

---

<div class="post-metadata">

**Author:** ![jondeuce](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jondeuce/32/16378_2.png) [@jondeuce](https://discourse.julialang.org/u/jondeuce)\
**Post date:** [April 6, 2021, 9:32pm UTC](https://discourse.julialang.org/t/softmax-and-large-numbers/58674/8 "2021-04-06T21:32:41Z")

</div>

This issue came up in a recent major rewriting of the CUDA internals by @denizyuret: [https://github.com/JuliaGPU/CUDA.jl/pull/523#issuecomment-753416384](https://github.com/JuliaGPU/CUDA.jl/pull/523#issuecomment-753416384)

As far as I understand, following this rewrite CUDA’s softmax should default to “accurate” arithmetic by default (via the `CUDNN_SOFTMAX_ACCURATE` flag), which should properly handle infinities. So it may also be an issue of which CUDA and/or NNlib and/or Julia version you are using.
