I still say congrats on that and more, but Julia 1.13 is not faster at everything (what follows mainly important for showing good benchmark numbers). I think I know why, likely mostly this overhead [see answer below]:
#= 298.8 ms =# precompile(Tuple{typeof(Main.main), Int64})
[it seems to be slower on 1.13, but hard to just try to eyeball it.]
and sometimes it goes up to:
#= 351.6 ms =# precompile(Tuple{typeof(Main.main), Int64})
and what is this new (shown in different color):
#= 17.3 ms =# precompile(Tuple{typeof(Base.should_use_main_entrypoint)}) # recompile
Background, I’m timing (can anyone confirm, it’s not just on my machine):
$ hyperfine 'julia +1.13 -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500'
Benchmark 1: julia +1.13 -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500
Time (mean ± σ): 1.538 s ± 0.030 s [User: 6.750 s, System: 0.160 s]
Range (min … max): 1.477 s … 1.576 s 10 runs
$ hyperfine 'julia +1.11.1 -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500'
Benchmark 1: julia +1.11.1 -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500
Time (mean ± σ): 1.210 s ± 0.059 s [User: 7.902 s, System: 0.113 s]
Range (min … max): 1.086 s … 1.272 s 10 runs
It's not about any extra flags such as -tauto.
Adding -O3 actually made is just slower possibly unlike for +1.11.1 that got tiny bit faster (or benchmark noise):
$ hyperfine 'julia +1.13 -O3 -- spectralnorm.julia-4.julia 5500'
Benchmark 1: julia +1.13 -O3 -- spectralnorm.julia-4.julia 5500
Time (mean ± σ): 3.572 s ± 0.064 s [User: 4.329 s, System: 0.127 s]
Range (min … max): 3.492 s … 3.699 s 10 runs
$ time julia +1.13 --trace-compile=stderr --trace-compile-timing --gc-sweep-always-full -- spectralnorm.julia-4.julia 55
#= 8.3 ms =# precompile(Tuple{typeof(Base.indexed_iterate), Tuple{QuoteNode, Expr}, Int64})
#= 6.5 ms =# precompile(Tuple{typeof(Base.indexed_iterate), Tuple{QuoteNode, Expr}, Int64, Int64})
#= 17.8 ms =# precompile(Tuple{typeof(Base.Threads._threadsfor), Expr, Expr, Symbol})
#= 15.1 ms =# precompile(Tuple{typeof(Base.Threads.default_func), Expr, Symbol, Expr})
#= 298.8 ms =# precompile(Tuple{typeof(Main.main), Int64})
#= 86.0 ms =# precompile(Tuple{Base.Threads.var"#threading_run##0#threading_run##1"{Main.var"#mul_by_f!##0#mul_by_f!##1"{Main.var"#mul_by_f!##2#mul_by_f!##3"{typeof(Main.A), Array{Float64, 1}, Array{Float64, 1}, Int64, Base.UnitRange{Int64}}}, Int64}})
#= 61.9 ms =# precompile(Tuple{Base.Threads.var"#threading_run##0#threading_run##1"{Main.var"#mul_by_f!##0#mul_by_f!##1"{Main.var"#mul_by_f!##2#mul_by_f!##3"{typeof(Main.At), Array{Float64, 1}, Array{Float64, 1}, Int64, Base.UnitRange{Int64}}}, Int64}})
#= 5.9 ms =# precompile(Tuple{typeof(Base.Ryu.writefixed), Float64, Int64})
1.274202174
#= 17.3 ms =# precompile(Tuple{typeof(Base.should_use_main_entrypoint)}) # recompile
real 0m0,931s
user 0m1,661s
sys 0m0,144s
$ time julia +1.11.1 --trace-compile=stderr -- spectralnorm.julia-4.julia 55
precompile(Tuple{typeof(Base.indexed_iterate), Tuple{QuoteNode, Expr}, Int64})
precompile(Tuple{typeof(Base.indexed_iterate), Tuple{QuoteNode, Expr}, Int64, Int64})
precompile(Tuple{typeof(Base.Threads._threadsfor), Expr, Expr, Symbol})
precompile(Tuple{typeof(Base.Threads.default_func), Expr, Symbol, Expr})
[I'm not sure how slow this is here or how to time it] precompile(Tuple{typeof(Main.main), Int64})
[similar text] precompile(Tuple{Base.Threads.var"#1#2"{Main.var"#2#threadsfor_fun#2"{Main.var"#2#threadsfor_fun#1#3"{typeof(Main.A), Array{Float64, 1}, Array{Float64, 1}, Int64, Base.UnitRange{Int64}}}, Int64}})
[similar text] precompile(Tuple{Base.Threads.var"#1#2"{Main.var"#2#threadsfor_fun#2"{Main.var"#2#threadsfor_fun#1#3"{typeof(Main.At), Array{Float64, 1}, Array{Float64, 1}, Int64, Base.UnitRange{Int64}}}, Int64}})
[only here] precompile(Tuple{typeof(Base.println), Base.TTY, String, Vararg{String}})
It's better on nightly, 1.14 will be awesome:
$ time julia +nightly -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500
1.274224153
real 0m0,910s
user 0m5,366s
sys 0m0,171s
vibe@ryksugan:~$ time julia +1.13 -tauto --cpu-target=ivybridge --math-mode=ieee -- spectralnorm.julia-4.julia 5500
1.274224153
real 0m2,170s
user 0m6,995s
sys 0m0,344s
Other still fixed overheads mostly startup itself and printing numbers…:
$ hyperfine 'julia +1.13 -e ""'
Benchmark 1: julia +1.13 -e ""
Time (mean ± σ): 242.0 ms ± 14.8 ms [User: 158.8 ms, System: 95.6 ms]
Range (min … max): 204.4 ms … 262.6 ms 12 runs
$ hyperfine 'julia +1.11.1 -e ""'
Benchmark 1: julia +1.11.1 -e ""
Time (mean ± σ): 195.7 ms ± 12.8 ms [User: 132.7 ms, System: 70.6 ms]
Range (min … max): 168.8 ms … 221.5 ms 15 runs
$ hyperfine 'julia +nightly -e ""'
Benchmark 1: julia +nightly -e ""
Time (mean ± σ): 236.8 ms ± 16.4 ms [User: 132.9 ms, System: 101.8 ms]
Range (min … max): 192.3 ms … 254.8 ms 12 runs
I thought I got precompilation for print and println fixed, but it needs to be more broad, floating-point numbers should be common...:
$ julia +1.13 --trace-compile=stderr --trace-compile-timing -e 'println(1.2)'
#= 24.4 ms =# precompile(Tuple{typeof(Base.println), Float64})
#= 59.6 ms =# precompile(Tuple{typeof(Base.print), Base.TTY, Float64, String})
[Base.Ryu.writefixed was precompiled in 1.11.1 but isn't in 1.13 or 1.10, and used in the benchmark for lack of precompiled float printing.]
$ julia +1.13 --trace-compile=stderr --trace-compile-timing -e 'println(Base.Ryu.writefixed(1.2, 9))'
#= 13.1 ms =# precompile(Tuple{typeof(Base.Ryu.writefixed), Float64, Int64})