# \#benchmark

**URL:** https://discourse.julialang.org/tag/benchmark/160.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [Implementing GPU Multi-Stream Capabilities to Dagger.jl](https://discourse.julialang.org/t/implementing-gpu-multi-stream-capabilities-to-dagger-jl/139013)

<div class="topic-metadata">

**Author:** [@camarav](https://discourse.julialang.org/u/camarav)\
**Replies:** 1\
**Last updated:** [August 27, 2026, 6:45pm UTC](https://discourse.julialang.org/t/implementing-gpu-multi-stream-capabilities-to-dagger-jl/139013 "2026-08-27T18:45:30Z")

</div>

Hello everyone! this post will be interesting especially for those who enjoy using 100% of what your computer has to offer, this summer I was part of GSoC and worked to implement multi-streams capabilities for GPUs in Da…

---

## [Benchmarking and threads](https://discourse.julialang.org/t/benchmarking-and-threads/138593)

<div class="topic-metadata">

**Author:** [@phma](https://discourse.julialang.org/u/phma)\
**Replies:** 9\
**Last updated:** [August 9, 2026, 9:14am UTC](https://discourse.julialang.org/t/benchmarking-and-threads/138593 "2026-08-09T09:14:02Z")

</div>

In the bench workspace, I have this function: function kaftorTime(textLen::Integer,keyLen::Integer) # in nanoseconds text=fill(0x69,textLen) key=fill(0x96,keyLen) trial=@benchmark kaftorEncrypt!($text,$key) me…

---

## [Allocations from scalar arguments in functions](https://discourse.julialang.org/t/allocations-from-scalar-arguments-in-functions/136927)

<div class="topic-metadata">

**Author:** [@Pablo\_Montes](https://discourse.julialang.org/u/Pablo_Montes)\
**Replies:** 3\
**Last updated:** [April 30, 2026, 2:56pm UTC](https://discourse.julialang.org/t/allocations-from-scalar-arguments-in-functions/136927 "2026-04-30T14:56:08Z")

</div>

Hello, I am trying to optimize a code for simulating fluids that is currently allocating massive amounts of memory. I’m focusing on some specific functions that are simple but need to be called repeatedly (thousands of t…

---

## [Variables in @belapsed benchmark](https://discourse.julialang.org/t/variables-in-belapsed-benchmark/134105)

<div class="topic-metadata">

**Author:** [@MrCapa](https://discourse.julialang.org/u/MrCapa)\
**Replies:** 2\
**Last updated:** [November 25, 2025, 5:45am UTC](https://discourse.julialang.org/t/variables-in-belapsed-benchmark/134105 "2025-11-25T05:45:03Z")

</div>

Hey there! Maybe anybody can help me? I don’t understand how variables work in @belapsed. My code: function test\_queue\_deque(n) @belapsed begin for i in 1:n popfirst!(dq) end end set…

---

## [Productive Scalable Distributed Task Scheduling Using an MPI-based Backend for Dagger](https://discourse.julialang.org/t/productive-scalable-distributed-task-scheduling-using-an-mpi-based-backend-for-dagger/131995)

<div class="topic-metadata">

**Author:** [@yanzin00](https://discourse.julialang.org/u/yanzin00)\
**Replies:** 4\
**Last updated:** [September 8, 2025, 4:45pm UTC](https://discourse.julialang.org/t/productive-scalable-distributed-task-scheduling-using-an-mpi-based-backend-for-dagger/131995 "2025-09-08T16:45:56Z")

</div>

Hello Julia Community, especially Dagger and HPC developers. I hope you are doing well this summer and have achieved your goals! I’m here to share: Dagger’s MPI Implementation status My Google Summer of Code at Julia c…

---

## [Discrepancy of ODE Sensitivity Analysis paper results with benchmarks](https://discourse.julialang.org/t/discrepancy-of-ode-sensitivity-analysis-paper-results-with-benchmarks/129853)

<div class="topic-metadata">

**Author:** [@mra](https://discourse.julialang.org/u/mra)\
**Replies:** 31\
**Last updated:** [June 18, 2025, 2:01pm UTC](https://discourse.julialang.org/t/discrepancy-of-ode-sensitivity-analysis-paper-results-with-benchmarks/129853 "2025-06-18T14:01:19Z")

</div>

I have tried to reproduce results from the paper “A Comparison of Automatic Differentiation and Continuous Sensitivity Analysis for Derivatives of Differential Equation Solutions” (https://arxiv.org/pdf/1812.01892, Fig. …

---

## [Yet another language benchmark](https://discourse.julialang.org/t/yet-another-language-benchmark/127924)

<div class="topic-metadata">

**Author:** [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Replies:** 9\
**Last updated:** [June 15, 2025, 12:13pm UTC](https://discourse.julialang.org/t/yet-another-language-benchmark/127924 "2025-06-15T12:13:00Z")

</div>

I like the visuals in this one. As usual, Julia compilation seems to be included in the results: P.S.: we don’t have a HPC category on discourse. Should it be created? The GPU and Julia at Scale categories could be su…

---

## [CI : benchmarking dependencies best practices?](https://discourse.julialang.org/t/ci-benchmarking-dependencies-best-practices/128402)

<div class="topic-metadata">

**Author:** [@s\_amap](https://discourse.julialang.org/u/s_amap)\
**Replies:** 3\
**Last updated:** [May 8, 2025, 12:12am UTC](https://discourse.julialang.org/t/ci-benchmarking-dependencies-best-practices/128402 "2025-05-08T00:12:12Z")

</div>

I’m working on several packages, with one or two packages with core features, and others depending on them, which aren’t really develop-centric, more so user applications. When adding features to the core packages, ther…

---

## [How to call blas getrf, getri properly ? i want to create a benchmark inverse matrix using gauss, crout, native julia inv(A) and BLAS direct](https://discourse.julialang.org/t/how-to-call-blas-getrf-getri-properly-i-want-to-create-a-benchmark-inverse-matrix-using-gauss-crout-native-julia-inv-a-and-blas-direct/128771)

<div class="topic-metadata">

**Author:** [@Roberto\_Bagyo](https://discourse.julialang.org/u/Roberto_Bagyo)\
**Replies:** 5\
**Last updated:** [May 7, 2025, 5:38pm UTC](https://discourse.julialang.org/t/how-to-call-blas-getrf-getri-properly-i-want-to-create-a-benchmark-inverse-matrix-using-gauss-crout-native-julia-inv-a-and-blas-direct/128771 "2025-05-07T17:38:06Z")

</div>

hello, i am new here. fortran based but just for fun actually. i tried to convert test\_fpu.f90 from polyhedron benchmark inverse matrix. i heard julia is multihreading and easier to code, then dgemm performance is highe…

---

## [Reducing the number of allocations in a function](https://discourse.julialang.org/t/reducing-the-number-of-allocations-in-a-function/126906)

<div class="topic-metadata">

**Author:** [@andrea\_bertozzi](https://discourse.julialang.org/u/andrea_bertozzi)\
**Replies:** 14\
**Last updated:** [March 14, 2025, 12:30pm UTC](https://discourse.julialang.org/t/reducing-the-number-of-allocations-in-a-function/126906 "2025-03-14T12:30:39Z")

</div>

Hello, I have this function. When running benchmark, I get 519 allocations. Does anyone know how I can reduce those? Thank you! using LinearAlgebra struct Settings rho::Float64 g\_earth::Vector{Float64} cd…

---

## [Benchmarking function compile time](https://discourse.julialang.org/t/benchmarking-function-compile-time/125013)

<div class="topic-metadata">

**Author:** [@AntonReinhard](https://discourse.julialang.org/u/AntonReinhard)\
**Replies:** 10\
**Last updated:** [January 22, 2025, 10:33am UTC](https://discourse.julialang.org/t/benchmarking-function-compile-time/125013 "2025-01-22T10:33:11Z")

</div>

I have some code that generates functions at runtime that vary greatly in size (10 to 100k lines of code). I’d like to benchmark the compilation of these functions. Is there any intended way to do this? Something like a …

---

## [Numpy.sort vs Julia sort](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421)

<div class="topic-metadata">

**Author:** [@gitboy16](https://discourse.julialang.org/u/gitboy16)\
**Replies:** 10\
**Last updated:** [December 4, 2024, 10:17pm UTC](https://discourse.julialang.org/t/numpy-sort-vs-julia-sort/123421 "2024-12-04T22:17:08Z")

</div>

I recently tried numpy.sort and it is really fast for a vector of 100 millions floats. It seems to be parallelized by default and/or using AVX-512. It is like 5 times faster on my system. If someone has a benchmark vs J…

---

## [Rule of thumb to estimate theoretical lower bound on benchmarks?](https://discourse.julialang.org/t/rule-of-thumb-to-estimate-theoretical-lower-bound-on-benchmarks/121509)

<div class="topic-metadata">

**Author:** [@Tetrakai](https://discourse.julialang.org/u/Tetrakai)\
**Replies:** 4\
**Last updated:** [October 25, 2024, 4:41pm UTC](https://discourse.julialang.org/t/rule-of-thumb-to-estimate-theoretical-lower-bound-on-benchmarks/121509 "2024-10-25T16:41:32Z")

</div>

I have a somewhat complex inner loop (I can share if needed but I am looking for general rules here) that benchmarks at 200-800 ns with no allocations. A range makes sense, it depends on which conditionals are randomly u…

---

## [Why does the speed of Julia in AOT compilation differ from UX4?](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657)

<div class="topic-metadata">

**Author:** [@aaoo](https://discourse.julialang.org/u/aaoo)\
**Replies:** 12\
**Last updated:** [September 21, 2024, 3:14am UTC](https://discourse.julialang.org/t/why-does-the-speed-of-julia-in-aot-compilation-differ-from-ux4/119657 "2024-09-21T03:14:26Z")

</div>

As the following picture shows, there are two different Julia versions in the speed comparison. The AOT compiled version reaches speeds as fast as Fortran, etc. However, the other version is slower and closer to Python. …

---

## [Get the maximum value by row in a DataFrame containing missing values](https://discourse.julialang.org/t/get-the-maximum-value-by-row-in-a-dataframe-containing-missing-values/118777)

<div class="topic-metadata">

**Author:** [@feanor12](https://discourse.julialang.org/u/feanor12)\
**Replies:** 10\
**Last updated:** [August 30, 2024, 5:37am UTC](https://discourse.julialang.org/t/get-the-maximum-value-by-row-in-a-dataframe-containing-missing-values/118777 "2024-08-30T05:37:15Z")

</div>

I want to aggregate data over rows, which works well if the row contains some values. When the data row is empty maximum throws an error. Maybe someone has an idea how to make this faster. My table has usually about r…

---

## [Ray Tracing in a week-end - Julia vs SIMD-optimized C++](https://discourse.julialang.org/t/ray-tracing-in-a-week-end-julia-vs-simd-optimized-c/72958)

<div class="topic-metadata">

**Author:** [@claforte](https://discourse.julialang.org/u/claforte)\
**Replies:** 81\
**Last updated:** [August 9, 2024, 6:36pm UTC](https://discourse.julialang.org/t/ray-tracing-in-a-week-end-julia-vs-simd-optimized-c/72958 "2024-08-09T18:36:10Z")

</div>

Hi folks, As an exercise to learn how to optimize Julia code for performance, I adapted from C++ to Julia a popular raytracing “book” (Ray Tracing In One Weekend by Peter Shirley) into a Pluto.jl notebook. Hopefully thi…

---

## [\[ANN\] AirspeedVelocity.jl - easily benchmark Julia packages over their lifetime](https://discourse.julialang.org/t/ann-airspeedvelocity-jl-easily-benchmark-julia-packages-over-their-lifetime/97221)

<div class="topic-metadata">

**Author:** [@MilesCranmer](https://discourse.julialang.org/u/MilesCranmer)\
**Replies:** 7\
**Last updated:** [June 19, 2024, 6:59am UTC](https://discourse.julialang.org/t/ann-airspeedvelocity-jl-easily-benchmark-julia-packages-over-their-lifetime/97221 "2024-06-19T06:59:32Z")

</div>

AirspeedVelocity.jl AirspeedVelocity.jl tries to make it easier to benchmark Julia packages over their lifetime. It is inspired by, and takes its name from, asv, (and aspires to one day have as nice a UI). Basically,…

---

## [Allocations and time of running the same program twice are orders of magnitude larger than running them separately](https://discourse.julialang.org/t/allocations-and-time-of-running-the-same-program-twice-are-orders-of-magnitude-larger-than-running-them-separately/114078)

<div class="topic-metadata">

**Author:** [@dieg0](https://discourse.julialang.org/u/dieg0)\
**Replies:** 11\
**Last updated:** [June 17, 2024, 12:58pm UTC](https://discourse.julialang.org/t/allocations-and-time-of-running-the-same-program-twice-are-orders-of-magnitude-larger-than-running-them-separately/114078 "2024-06-17T12:58:52Z")

</div>

I am relatively new to Julia. I am building a code to optimize some objective function that is called many times. The most important intermediate step is computing the following function, called recover\_δ! which is a fix…

---

## [Why is flux model slower than python?](https://discourse.julialang.org/t/why-is-flux-model-slower-than-python/54154)

<div class="topic-metadata">

**Author:** [@mdsa3d](https://discourse.julialang.org/u/mdsa3d)\
**Replies:** 6\
**Last updated:** [May 11, 2024, 2:50pm UTC](https://discourse.julialang.org/t/why-is-flux-model-slower-than-python/54154 "2024-05-11T14:50:44Z")

</div>

I ran a benchmark test to observe the performance gain using flux. But, rather I observed that the flux runtine is longer than python equivalent. However, the julia is supposed to be faster. May I know, what am I doing …

---

## [BenchmarkTools needs a maintainer](https://discourse.julialang.org/t/benchmarktools-needs-a-maintainer/111178)

<div class="topic-metadata">

**Author:** [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Replies:** 10\
**Last updated:** [May 2, 2024, 3:42pm UTC](https://discourse.julialang.org/t/benchmarktools-needs-a-maintainer/111178 "2024-05-02T15:42:16Z")

</div>

Hi everyone, I have been maintaining BenchmarkTools.jl for about a year. I didn’t particularly want to do it, nor did the person in charge before me, but we did our best to keep it afloat. Now I am stepping down, and lo…

---

## [Performance discrepancy with multiple dispatch](https://discourse.julialang.org/t/performance-discrepancy-with-multiple-dispatch/113556)

<div class="topic-metadata">

**Author:** [@Albert\_de\_montserrat](https://discourse.julialang.org/u/Albert_de_montserrat)\
**Replies:** 4\
**Last updated:** [April 27, 2024, 11:07pm UTC](https://discourse.julialang.org/t/performance-discrepancy-with-multiple-dispatch/113556 "2024-04-27T23:07:26Z")

</div>

I am having problems trying to understand the difference in performance between the two foo methods described below. using Chairmarks abstract type AbstractTrait end struct ConcreteTrait \<: AbstractTrait end foo(::Arra…

---

## [Difference in microbenchmark result, Chairmarks.jl vs BenchmarkTools.jl](https://discourse.julialang.org/t/difference-in-microbenchmark-result-chairmarks-jl-vs-benchmarktools-jl/111656)

<div class="topic-metadata">

**Author:** [@TheLateKronos](https://discourse.julialang.org/u/TheLateKronos)\
**Replies:** 8\
**Last updated:** [March 20, 2024, 11:03am UTC](https://discourse.julialang.org/t/difference-in-microbenchmark-result-chairmarks-jl-vs-benchmarktools-jl/111656 "2024-03-20T11:03:40Z")

</div>

Hi all. I am using BenchmarkTools.jl and Chairmarks.jl to compare the performance of the following (hopefully) mathematically equivalent expressions, to see if they are computationally equivalent. The results are as foll…

---

## [Benchmarking surprises with type assertions](https://discourse.julialang.org/t/benchmarking-surprises-with-type-assertions/111660)

<div class="topic-metadata">

**Author:** [@sijo](https://discourse.julialang.org/u/sijo)\
**Replies:** 10\
**Last updated:** [March 18, 2024, 1:48pm UTC](https://discourse.julialang.org/t/benchmarking-surprises-with-type-assertions/111660 "2024-03-18T13:48:23Z")

</div>

While looking at another post I got surprised by the following differences when using type assertions: using Chairmarks tup = (10, 100im, 1.0, 2//3, 3.0im, 2//3+im) idx = 5 type = typeof(tup\[idx\]) function test1(tup, …

---

## [Chairmarks.jl](https://discourse.julialang.org/t/chairmarks-jl/111096)

<div class="topic-metadata">

**Author:** [@Lilith](https://discourse.julialang.org/u/Lilith)\
**Replies:** 82\
**Last updated:** [March 12, 2024, 7:45pm UTC](https://discourse.julialang.org/t/chairmarks-jl/111096 "2024-03-12T19:45:39Z")

</div>

Hello folks! I’m announcing Chairmarks.jl, version 1.0 (docs). It’s a benchmarking package that aims to be hundreds of times faster than BenchmarkTools without compromising on precision. Usage is pretty similar to Ben…

---

## [Benchmarking function in local scope](https://discourse.julialang.org/t/benchmarking-function-in-local-scope/110228)

<div class="topic-metadata">

**Author:** [@leespen1](https://discourse.julialang.org/u/leespen1)\
**Replies:** 4\
**Last updated:** [February 16, 2024, 7:48pm UTC](https://discourse.julialang.org/t/benchmarking-function-in-local-scope/110228 "2024-02-16T19:48:56Z")

</div>

I have a function that I would like to benchmark. To that end, I would like to time the function for several different argument values, and plot how the runtime changes with the argument values. I can do this using the @…

---

## [Cuda (Julia vs C++)](https://discourse.julialang.org/t/cuda-julia-vs-c/109865)

<div class="topic-metadata">

**Author:** [@Enlil50](https://discourse.julialang.org/u/Enlil50)\
**Replies:** 4\
**Last updated:** [February 12, 2024, 12:03pm UTC](https://discourse.julialang.org/t/cuda-julia-vs-c/109865 "2024-02-12T12:03:44Z")

</div>

Are there any benchmarks between C++ cuda and Cuda.jl?

---

## [Billion-row benchmark?](https://discourse.julialang.org/t/billion-row-benchmark/108417)

<div class="topic-metadata">

**Author:** [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Replies:** 1\
**Last updated:** [January 5, 2024, 10:45pm UTC](https://discourse.julialang.org/t/billion-row-benchmark/108417 "2024-01-05T22:45:03Z")

</div>

You’re correct - I had no intention to compare with Python or start a debate - the extract was to highlight that he either didnt know of Julia or didnt consider it (despite aggrievances with Python). In so far as the be…

---

## [Trivial port from C++/Rust to Julia](https://discourse.julialang.org/t/trivial-port-from-c-rust-to-julia/105787)

<div class="topic-metadata">

**Author:** [@feribg](https://discourse.julialang.org/u/feribg)\
**Replies:** 51\
**Last updated:** [December 28, 2023, 5:21am UTC](https://discourse.julialang.org/t/trivial-port-from-c-rust-to-julia/105787 "2023-12-28T05:21:30Z")

</div>

I’m trying to port a very trivial MC simulation from C++ / Rust to Julia. Here are the 3 implementations: C++: (clang OSX) https://pastebin.com/HWEUxa6W Rust: https://pastebin.com/7d31ynJE Julia: https://pastebin.com…

---

## [Multi-threaded benchmarking is slow](https://discourse.julialang.org/t/multi-threaded-benchmarking-is-slow/106386)

<div class="topic-metadata">

**Author:** [@filchristou](https://discourse.julialang.org/u/filchristou)\
**Replies:** 12\
**Last updated:** [November 21, 2023, 7:43pm UTC](https://discourse.julialang.org/t/multi-threaded-benchmarking-is-slow/106386 "2023-11-21T19:43:32Z")

</div>

Motivation Running many benchmarks can be slow. Since almost all of them are independent, they can theoretically run easily on parallel. Problem It turns out running the benchmarks in parallel takes more or less equal …

---

## [How to asychronously build a benchmark group](https://discourse.julialang.org/t/how-to-asychronously-build-a-benchmark-group/106335)

<div class="topic-metadata">

**Author:** [@filchristou](https://discourse.julialang.org/u/filchristou)\
**Replies:** 3\
**Last updated:** [November 17, 2023, 8:35pm UTC](https://discourse.julialang.org/t/how-to-asychronously-build-a-benchmark-group/106335 "2023-11-17T20:35:04Z")

</div>

Can I build several inter-depending benchmarks from a serial procedure in “one go” using BenchmarkTools.jl ? Say I have the following functionality laid out in a function function mymutatingfunc!(x) mutfunc1!(x) …

[Next page](https://discourse.julialang.org/tag/benchmark/160.md?match_all_tags=true&page=1&tags%5B%5D=benchmark)
