# Package/workflow for combining BenchmarkTools.jl and HyphothesisTests.jl

**URL:** <https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252>\
**Category:** General Usage\
**Created:** [September 26, 2023, 10:09am UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252 "2023-09-26T10:09:13Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [September 26, 2023, 10:09am UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/1 "2023-09-26T10:09:13Z")

</div>

It seems to me that we should have the technology to determine if a change in the performance of a function is statistically significant, possibly with some user specified assumptions.

Has anyone made a workflow or little package for this?

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [September 26, 2023, 11:59am UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/2 "2023-09-26T11:59:09Z")

</div>

Note that benchmarking data is non-iid and non-normal.  
So something like a t-test would be inappropriate.

So you likely want two-sample Kolmogorov-Smirnov or Anderson-Darling, but I would need to double check if either have normality assumptions.

---

<div class="post-metadata">

**Author:** ![aplavin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aplavin/32/222056_2.png) [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Post date:** [September 26, 2023, 12:14pm UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/3 "2023-09-26T12:14:43Z")

</div>

Permutation test should be most robust statistically. Also, to mitigate the effect of always varying computer load, running two functions in an interleaved way would probably help. Like, 1000x run `f()`, 1000x run `g()`, repeat these 100x.

---

<div class="post-metadata">

**Author:** ![tbeason](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tbeason/32/15898_2.png) [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Post date:** [September 26, 2023, 5:55pm UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/4 "2023-09-26T17:55:15Z")

</div>

Is the full distribution of run times kept or just the histogram? If you have all times, might as well just use it as a “bootstrap” distribution. No need to rely on distributional assumptions if you have the distribution.

---

<div class="post-metadata">

**Author:** ![vchuravy](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/vchuravy/32/8_2.png) [@vchuravy](https://discourse.julialang.org/u/vchuravy)\
**Post date:** [September 26, 2023, 8:16pm UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/5 "2023-09-26T20:16:57Z")

</div>

We have all data points available. Do you have a reference for a bootstrap method?

The problem is described as future work in [[1608.04295] Robust benchmarking in noisy environments](https://arxiv.org/abs/1608.04295) IIRC.

Having a more reliable comparison between two benchmark runs would be very powerful.

---

<div class="post-metadata">

**Author:** ![Christopher\_Fisher](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/christopher_fisher/32/26132_2.png) [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Post date:** [September 26, 2023, 9:04pm UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/6 "2023-09-26T21:04:19Z")

</div>

How many repetitions do you normally have of a given benchmark? The documentation states:

- `samples`: The number of samples to take. Execution will end if this many samples have been collected. Defaults to `BenchmarkTools.DEFAULT_PARAMETERS.samples = 10000`.

If that is a typical sample size, one concern is that the test would be sensitive to very small discrepancies. Another thing to consider the number of flagged benchmarks. In a large suite of benchmarks, you may get a large number of spurious flags (approximately 5% under most assumptions). I wonder if an effect size statistic might be a better way to evaluate differences?

---

<div class="post-metadata">

**Author:** ![gdalle](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/gdalle/32/27854_2.png) [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Post date:** [September 26, 2023, 10:26pm UTC](https://discourse.julialang.org/t/package-workflow-for-combining-benchmarktools-jl-and-hyphothesistests-jl/104252/7 "2023-09-26T22:26:03Z")

</div>

Also related: [Making @benchmark outputs statistically meaningful, and actionable](https://discourse.julialang.org/t/making-benchmark-outputs-statistically-meaningful-and-actionable/104256/)

And the corresponding issue on GitHub: [use legitimate non-iid hypothesis testing · Issue #74 · JuliaCI/BenchmarkTools.jl · GitHub](https://github.com/JuliaCI/BenchmarkTools.jl/issues/74)

Don’t mind me, I’m just doing some sweet sweet cross-referencing
