# Speeding up antialiased non-linearity

**URL:** <https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456>\
**Category:** Machine Learning\
**Created:** [June 5, 2026, 9:52am UTC](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456 "2026-06-05T09:52:43Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![fps](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fps/32/216317_2.png) [@fps](https://discourse.julialang.org/u/fps)\
**Post date:** [June 5, 2026, 9:52am UTC](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456/1 "2026-06-05T09:52:43Z")

</div>

Hi,

I’m playing around with some guitar amplifier modeling and in the course of that I implemented the anti-derivative anti-aliased non-linearity of “Note on Alias Suppression in Digital Distortion”, Martin Vicanek in a naive form:

```julia
function dist_aa2(x)
  x0 = x[3:end,:,:]
  x1 = x[2:(end-1),:,:]
  x2 = x[1:(end-2),:,:]

  F1 = sqrt.(1 .+ x1.^2)
  F12 = sqrt.(1 .+ ((x0 + x1)./2).^2)
  F23 = sqrt.(1 .+ ((x1 + x2)./2).^2)

  ((x0 .+ 3 .* x1) ./ (F12 .+ F1) .+ (x2 .+ 3 .* x1) ./ (F23 .+ F1)) ./ 4
end

```

While it works in my experiments (i.e. training converges and produces nice results) it seems to be quite slow in the current form and wonder if there is an obvious way to speed it up. I’m using Flux with CUDA.jl/cuDNN.jl and use it as activation by using a `Flux.Chain(Flux.Conv(#= 1D convolution without activation here =#), dist_aa2)` as a single layer.

Thanks!

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [June 5, 2026, 12:13pm UTC](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456/2 "2026-06-05T12:13:30Z")

</div>

You are allocating lots of temporary arrays. At least on a CPU, I would just write a single scalar function and then broadcast it over views:

```julia-auto
function dist_aa2(x0, x1, x2)
  F1 = sqrt(1 + x1^2)
  F12 = sqrt(1 + ((x0 + x1)/2)^2)
  F23 = sqrt(1 + ((x1 + x2)/2)^2)
  return ((x0 + 3 * x1) / (F12 + F1) + (x2 + 3 * x1) / (F23 + F1)) * (1//4)
end

dist_aa2(x) = @views dist_aa2.(x[3:end,:,:], x[2:(end-1),:,:], x[1:(end-2),:,:])

```

(For `x = rand(100,100,100)` on my CPU, though, this is only about a 40% speedup.)

---

<div class="post-metadata">

**Author:** ![fps](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fps/32/216317_2.png) [@fps](https://discourse.julialang.org/u/fps)\
**Post date:** [June 10, 2026, 7:30am UTC](https://discourse.julialang.org/t/speeding-up-antialiased-non-linearity/137456/3 "2026-06-10T07:30:54Z")

</div>

It seems to help on the GPU as well! Thanks for the suggestion!
