# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=30

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 31

---

## [Unexpected memory allocation](https://discourse.julialang.org/t/unexpected-memory-allocation/109094)

<div class="topic-metadata">

**Author:** [@JonasKoziorek](https://discourse.julialang.org/u/JonasKoziorek)\
**Replies:** 2\
**Last updated:** [January 22, 2024, 12:25pm UTC](https://discourse.julialang.org/t/unexpected-memory-allocation/109094 "2024-01-22T12:25:57Z")

</div>

Hello, I have the following code: using ForwardDiff function logistic(x, p) return p\*x\*(1-x) end function nth\_composite(func, x0, p, n) x = x0 for \_ in 1:n x = func(x, p) end return x end …

---

## [PrecompileTools and --trace-compile](https://discourse.julialang.org/t/precompiletools-and-trace-compile/109067)

<div class="topic-metadata">

**Author:** [@mchitre](https://discourse.julialang.org/u/mchitre)\
**Replies:** 3\
**Last updated:** [January 21, 2024, 3:26pm UTC](https://discourse.julialang.org/t/precompiletools-and-trace-compile/109067 "2024-01-21T15:26:22Z")

</div>

I have a simple package Test1: module Test1 f1(x) = x + 1 f2(x) = f1(2x) f3(x) = f1(3x) using PrecompileTools @compile\_workload begin f2(1.2) end end # module I load and precompile the package in an environment. I…

---

## [CSV vs DelimitedFiles vs Numpy](https://discourse.julialang.org/t/csv-vs-delimitedfiles-vs-numpy/108963)

<div class="topic-metadata">

**Author:** [@Matt\_jl](https://discourse.julialang.org/u/Matt_jl)\
**Replies:** 15\
**Last updated:** [January 20, 2024, 4:38pm UTC](https://discourse.julialang.org/t/csv-vs-delimitedfiles-vs-numpy/108963 "2024-01-20T16:38:34Z")

</div>

I’ve been working on a project where I need to read specific rows and columns from a data file. To determine the most efficient approach, I conducted benchmarks using CSV.jl, DelimitedFiles.jl, and Numpy in Python. The r…

---

## [abs(::Int32) slower than abs(::Int)](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011)

<div class="topic-metadata">

**Author:** [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Replies:** 7\
**Last updated:** [January 19, 2024, 8:18pm UTC](https://discourse.julialang.org/t/abs-int32-slower-than-abs-int/109011 "2024-01-19T20:18:30Z")

</div>

Hi, I have a need for a lot of large integer vectors, so have attempted to make use of Int32 rather than Int to minimise the memory footprint. As I need to ensure the values are positive, I call abs(i) on each element d…

---

## [How to improve performances of this multiplication -\> sum](https://discourse.julialang.org/t/how-to-improve-performances-of-this-multiplication-sum/108955)

<div class="topic-metadata">

**Author:** [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Replies:** 6\
**Last updated:** [January 18, 2024, 1:48pm UTC](https://discourse.julialang.org/t/how-to-improve-performances-of-this-multiplication-sum/108955 "2024-01-18T13:48:08Z")

</div>

I am trying hard to improve the performances of this function that just sum a serie of multiplications, but where all the data is given in terms of indices: This is my data: using BenchmarkTools, LoopVectorization x …

---

## [Hand written loop slower than bitVector broadcast](https://discourse.julialang.org/t/hand-written-loop-slower-than-bitvector-broadcast/108881)

<div class="topic-metadata">

**Author:** [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Replies:** 15\
**Last updated:** [January 18, 2024, 2:56am UTC](https://discourse.julialang.org/t/hand-written-loop-slower-than-bitvector-broadcast/108881 "2024-01-18T02:56:42Z")

</div>

Hi, MWE below I have a use case for intersection of lists of sorted integer vectors with a vector of sorted integers i.e. intersect!(vec :: Vector{Int},vectorOfVectors::Vector{Vector{Int}}) because the vectorOfVecto…

---

## [Improving speed with iterative sums and functions within functions](https://discourse.julialang.org/t/improving-speed-with-iterative-sums-and-functions-within-functions/108927)

<div class="topic-metadata">

**Author:** [@clairevalva](https://discourse.julialang.org/u/clairevalva)\
**Replies:** 4\
**Last updated:** [January 17, 2024, 9:56pm UTC](https://discourse.julialang.org/t/improving-speed-with-iterative-sums-and-functions-within-functions/108927 "2024-01-17T21:56:44Z")

</div>

Hi all — I’m having some problems making this bit of code run at a reasonable speed. In essence, I have functions defined within functions, where the innermost function evaluates a sum. Then, I end up combining functions…

---

## [From MKL back to OpenBLAS](https://discourse.julialang.org/t/from-mkl-back-to-openblas/37722)

<div class="topic-metadata">

**Author:** [@martincornejo](https://discourse.julialang.org/u/martincornejo)\
**Replies:** 15\
**Last updated:** [January 17, 2024, 2:04am UTC](https://discourse.julialang.org/t/from-mkl-back-to-openblas/37722 "2024-01-17T02:04:06Z")

</div>

After building MKL on Julia. julia\>\]add https://github.com/JuliaComputing/MKL.jl julia\>\] build MKL Is there a way to set OpenBLAS back?

---

## [Generated LLVM code causes register spills on x86\_64](https://discourse.julialang.org/t/generated-llvm-code-causes-register-spills-on-x86-64/108144)

<div class="topic-metadata">

**Author:** [@akrishnamoorthy](https://discourse.julialang.org/u/akrishnamoorthy)\
**Replies:** 11\
**Last updated:** [January 16, 2024, 8:30pm UTC](https://discourse.julialang.org/t/generated-llvm-code-causes-register-spills-on-x86-64/108144 "2024-01-16T20:30:55Z")

</div>

This question concerns circshift implemented here, which is reproduced below. function circshift(x::Tuple, shift::Integer) @inline j = mod1(shift, length(x)) ntuple(k -\> \_\_safe\_getindex(x, k-j+ifelse(k\>j,0,l…

---

## [Excessive Memory Allocation?](https://discourse.julialang.org/t/excessive-memory-allocation/108841)

<div class="topic-metadata">

**Author:** [@freestatelabs](https://discourse.julialang.org/u/freestatelabs)\
**Replies:** 6\
**Last updated:** [January 16, 2024, 9:03am UTC](https://discourse.julialang.org/t/excessive-memory-allocation/108841 "2024-01-16T09:03:57Z")

</div>

I have a program I’m writing that reads in two Matrices, performs a bit of linear algebra in a pair of nested loops, then outputs a new Matrix. Because eventually it needs to operate on very large input Matrices, I’ve be…

---

## [Multiplication and increment between elements whose index is stored in vectors without allocation: possible?](https://discourse.julialang.org/t/multiplication-and-increment-between-elements-whose-index-is-stored-in-vectors-without-allocation-possible/108819)

<div class="topic-metadata">

**Author:** [@sylvaticus](https://discourse.julialang.org/u/sylvaticus)\
**Replies:** 5\
**Last updated:** [January 15, 2024, 6:25pm UTC](https://discourse.julialang.org/t/multiplication-and-increment-between-elements-whose-index-is-stored-in-vectors-without-allocation-possible/108819 "2024-01-15T18:25:26Z")

</div>

Hello, I have 3 arrays (x,y,w) where I want elementwise y += x \* w but the specific element indexes are stored in 3 separate vectors of SArrays, x\_ids, y\_ids and w\_ids. The problem is that the operation hugely allocates…

---

## [Fastest \`sqrt\` and \`log\` with negative check](https://discourse.julialang.org/t/fastest-sqrt-and-log-with-negative-check/107575)

<div class="topic-metadata">

**Author:** [@DanDoe](https://discourse.julialang.org/u/DanDoe)\
**Replies:** 17\
**Last updated:** [January 14, 2024, 12:23pm UTC](https://discourse.julialang.org/t/fastest-sqrt-and-log-with-negative-check/107575 "2024-01-14T12:23:15Z")

</div>

NaNMath.jl implements special functions such as sqrt and log which return NaN if called with negative argument. The implementation of sqrt is given by sqrt(x::T) where {T\<:AbstractFloat} = x \< 0.0 ? T(NaN) : Base.sqrt(…

---

## [Efficiency of parsing ASCII vs. Unicode](https://discourse.julialang.org/t/efficiency-of-parsing-ascii-vs-unicode/108727)

<div class="topic-metadata">

**Author:** [@greatpet](https://discourse.julialang.org/u/greatpet)\
**Replies:** 10\
**Last updated:** [January 12, 2024, 5:42pm UTC](https://discourse.julialang.org/t/efficiency-of-parsing-ascii-vs-unicode/108727 "2024-01-12T17:42:51Z")

</div>

What about performance implications for parsing ASCII text files? Consider the following function which checks if a file has an equal number of left brackets and right brackets: function file\_has\_balanced\_brackets(file)…

---

## [Union splitting not working in Julia 1.10?](https://discourse.julialang.org/t/union-splitting-not-working-in-julia-1-10/108710)

<div class="topic-metadata">

**Author:** [@PeterSimon](https://discourse.julialang.org/u/PeterSimon)\
**Replies:** 7\
**Last updated:** [January 12, 2024, 4:20pm UTC](https://discourse.julialang.org/t/union-splitting-not-working-in-julia-1-10/108710 "2024-01-12T16:20:51Z")

</div>

I’m getting the following results under 1.10: julia\> using BenchmarkTools julia\> x = rand(10\_000); julia\> function badsum(x) s = 0 for t in x s += t end …

---

## [Reduce allocations when extracting values from NamedTuples](https://discourse.julialang.org/t/reduce-allocations-when-extracting-values-from-namedtuples/108570)

<div class="topic-metadata">

**Author:** [@Christopher\_Fisher](https://discourse.julialang.org/u/Christopher_Fisher)\
**Replies:** 12\
**Last updated:** [January 12, 2024, 1:18pm UTC](https://discourse.julialang.org/t/reduce-allocations-when-extracting-values-from-namedtuples/108570 "2024-01-12T13:18:51Z")

</div>

Hello, I have a situation where I need to extract values from NamedTuples which are embedded within structs. Unfortunately, there are a lot of allocations, and it is not type stable. I have created a MWE below. mutable…

---

## [How to return all variable values of a non-simplified system of equations from a structurally-simplified solution output](https://discourse.julialang.org/t/how-to-return-all-variable-values-of-a-non-simplified-system-of-equations-from-a-structurally-simplified-solution-output/107915)

<div class="topic-metadata">

**Author:** [@chris-hampel-CA](https://discourse.julialang.org/u/chris-hampel-CA)\
**Replies:** 13\
**Last updated:** [January 11, 2024, 5:46pm UTC](https://discourse.julialang.org/t/how-to-return-all-variable-values-of-a-non-simplified-system-of-equations-from-a-structurally-simplified-solution-output/107915 "2024-01-11T17:46:00Z")

</div>

Hello, I am using MTK to solve a system of algebraic nonlinear equations. Using structural\_simplify() helps make these types of problems much easier to solve. However, I have noticed an inefficiency when trying to retu…

---

## [Using and understanding multi-threading](https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682)

<div class="topic-metadata">

**Author:** [@Andrea\_Vigliotti](https://discourse.julialang.org/u/Andrea_Vigliotti)\
**Replies:** 1\
**Last updated:** [January 11, 2024, 2:01pm UTC](https://discourse.julialang.org/t/using-and-understanding-multi-threading/108682 "2024-01-11T14:01:30Z")

</div>

so, I am trying to understand how multi threading works, I started julia as julia -t 16 on a 32 cores node where I reserved 8, then I run the example from the documentation here these are the lines function sum\_single(a…

---

## [Using Polyester.jl with ChunkSplitters.jl](https://discourse.julialang.org/t/using-polyester-jl-with-chunksplitters-jl/108528)

<div class="topic-metadata">

**Author:** [@Bruno\_Amorim](https://discourse.julialang.org/u/Bruno_Amorim)\
**Replies:** 4\
**Last updated:** [January 11, 2024, 1:33pm UTC](https://discourse.julialang.org/t/using-polyester-jl-with-chunksplitters-jl/108528 "2024-01-11T13:33:46Z")

</div>

I am trying to paralelize a function, where I have some quantity a that I split in chunks and paralelize over the differente chunks. After that I combine the results obtained from all chunks. An example of this would be …

---

## [Substantial increase in time in copying an instantiated broadcasted object vs a non-instantiated one](https://discourse.julialang.org/t/substantial-increase-in-time-in-copying-an-instantiated-broadcasted-object-vs-a-non-instantiated-one/108673)

<div class="topic-metadata">

**Author:** [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Replies:** 2\
**Last updated:** [January 11, 2024, 11:49am UTC](https://discourse.julialang.org/t/substantial-increase-in-time-in-copying-an-instantiated-broadcasted-object-vs-a-non-instantiated-one/108673 "2024-01-11T11:49:58Z")

</div>

I’m trying to understand the difference in performance here: julia\> using LinearAlgebra julia\> n=40; U = UpperTriangular(rand(n,n)); C = similar(U); B = Broadcast.broadcasted(\*, 2.0, U); julia\> @btime copyto!($C, $B);…

---

## [Setting up worker local buffers when using Distributed.jl](https://discourse.julialang.org/t/setting-up-worker-local-buffers-when-using-distributed-jl/108671)

<div class="topic-metadata">

**Author:** [@RomeoV](https://discourse.julialang.org/u/RomeoV)\
**Replies:** 2\
**Last updated:** [January 11, 2024, 11:20am UTC](https://discourse.julialang.org/t/setting-up-worker-local-buffers-when-using-distributed-jl/108671 "2024-01-11T11:20:10Z")

</div>

Hello all, I am in the process of incrementally optimizing an algorithm, having focused first on single-core performance, then threading, and now want to scale to Distributed workers. One key element of my optimization…

---

## [Poor Distributed performance for independent linear algebra operators](https://discourse.julialang.org/t/poor-distributed-performance-for-independent-linear-algebra-operators/108480)

<div class="topic-metadata">

**Author:** [@hshackle](https://discourse.julialang.org/u/hshackle)\
**Replies:** 9\
**Last updated:** [January 10, 2024, 9:29pm UTC](https://discourse.julialang.org/t/poor-distributed-performance-for-independent-linear-algebra-operators/108480 "2024-01-10T21:29:05Z")

</div>

I’m trying to run some embarrasingly parallel code using Distributed but am seeing poor scaling when scaling up the number of cores. The actual code is rather complicated but for the most part just involves a lot of line…

---

## [Struct and memory allocation](https://discourse.julialang.org/t/struct-and-memory-allocation/108632)

<div class="topic-metadata">

**Author:** [@fdekerme](https://discourse.julialang.org/u/fdekerme)\
**Replies:** 4\
**Last updated:** [January 10, 2024, 4:29pm UTC](https://discourse.julialang.org/t/struct-and-memory-allocation/108632 "2024-01-10T16:29:07Z")

</div>

Hello ! :slight\_smile: I don’t understand why there is memory allocation when I acces to a parameters of my immutable struct ? julia\> using BenchmarkTools julia\> struct S1 a::NTuple{3,Int64} b::Float64 …

---

## [PrecompileTools.@compile\_workload does not seem to work at all - what am I doing wrong?](https://discourse.julialang.org/t/precompiletools-compile-workload-does-not-seem-to-work-at-all-what-am-i-doing-wrong/108481)

<div class="topic-metadata">

**Author:** [@schlichtanders](https://discourse.julialang.org/u/schlichtanders)\
**Replies:** 4\
**Last updated:** [January 9, 2024, 3:24pm UTC](https://discourse.julialang.org/t/precompiletools-compile-workload-does-not-seem-to-work-at-all-what-am-i-doing-wrong/108481 "2024-01-09T15:24:38Z")

</div>

Hi there, I am testing PrecompileTools, and I just don’t get it to work - despite precompilation seems to work (check by MethodAnalysis.methodinstances), julia just recompiles again and again und process restart… module…

---

## [StaticArrays forces recompilation of JLLs on Windows with Julia 1.8.1](https://discourse.julialang.org/t/staticarrays-forces-recompilation-of-jlls-on-windows-with-julia-1-8-1/87725)

<div class="topic-metadata">

**Author:** [@mkitti](https://discourse.julialang.org/u/mkitti)\
**Replies:** 18\
**Last updated:** [January 9, 2024, 12:59pm UTC](https://discourse.julialang.org/t/staticarrays-forces-recompilation-of-jlls-on-windows-with-julia-1-8-1/87725 "2024-01-09T12:59:40Z")

</div>

Eric Davies, @longemen3000, I and others encountered an unexpected result after looking at @time\_imports and discussed it on Slack. Something was forcing JLL packages to recompile a lot of code. We narrowed it down to th…

---

## [How to optimize struct for speed?](https://discourse.julialang.org/t/how-to-optimize-struct-for-speed/108518)

<div class="topic-metadata">

**Author:** [@Alseidon](https://discourse.julialang.org/u/Alseidon)\
**Replies:** 4\
**Last updated:** [January 9, 2024, 8:53am UTC](https://discourse.julialang.org/t/how-to-optimize-struct-for-speed/108518 "2024-01-09T08:53:05Z")

</div>

I am defining a new struct for a project, which I need to be fast. This struct only contains one parameter, an MVector (from the StaticArrays.jl package). However, the operations on the struct are 4x slower than those on…

---

## [Tower hanoi in julia vs java - recursive function](https://discourse.julialang.org/t/tower-hanoi-in-julia-vs-java-recursive-function/108399)

<div class="topic-metadata">

**Author:** [@fab6](https://discourse.julialang.org/u/fab6)\
**Replies:** 14\
**Last updated:** [January 7, 2024, 3:58pm UTC](https://discourse.julialang.org/t/tower-hanoi-in-julia-vs-java-recursive-function/108399 "2024-01-07T15:58:33Z")

</div>

Hi, I am comparing a simple recursive function for the tower of hanoi in julia with java. This is the java code: class hanoi { static void towerOfHanoi(int n, char source, char target, char helper) { i…

---

## [Billion-row benchmark?](https://discourse.julialang.org/t/billion-row-benchmark/108417)

<div class="topic-metadata">

**Author:** [@djholiver](https://discourse.julialang.org/u/djholiver)\
**Replies:** 1\
**Last updated:** [January 5, 2024, 10:45pm UTC](https://discourse.julialang.org/t/billion-row-benchmark/108417 "2024-01-05T22:45:03Z")

</div>

You’re correct - I had no intention to compare with Python or start a debate - the extract was to highlight that he either didnt know of Julia or didnt consider it (despite aggrievances with Python). In so far as the be…

---

## [Group profiling flame graph tiles](https://discourse.julialang.org/t/group-profiling-flame-graph-tiles/108384)

<div class="topic-metadata">

**Author:** [@gdalle](https://discourse.julialang.org/u/gdalle)\
**Replies:** 4\
**Last updated:** [January 5, 2024, 9:17pm UTC](https://discourse.julialang.org/t/group-profiling-flame-graph-tiles/108384 "2024-01-05T21:17:01Z")

</div>

An integration + optimization code I’m working on generates flame graphs like this one, which are very hard to parse. Indeed, the same underlying function is called by several different lines of the code, so even if I sp…

---

## [Equivalent Fortran function is about 4.6 times faster](https://discourse.julialang.org/t/equivalent-fortran-function-is-about-4-6-times-faster/108344)

<div class="topic-metadata">

**Author:** [@jabru](https://discourse.julialang.org/u/jabru)\
**Replies:** 17\
**Last updated:** [January 5, 2024, 8:02am UTC](https://discourse.julialang.org/t/equivalent-fortran-function-is-about-4-6-times-faster/108344 "2024-01-05T08:02:08Z")

</div>

Hey there, for my current project, I wrote a function in plane Julia. It reads #Kappa squared or the square of the inverse debye length Debyelengthinnversesquared(is) = e^2\*beta\*2\*is\*n\_a/(epsilon\*epsilon\_r) #Kappa in …

---

## [Type stability with type as argument and permutedims](https://discourse.julialang.org/t/type-stability-with-type-as-argument-and-permutedims/108335)

<div class="topic-metadata">

**Author:** [@snowgum](https://discourse.julialang.org/u/snowgum)\
**Replies:** 3\
**Last updated:** [January 4, 2024, 12:08pm UTC](https://discourse.julialang.org/t/type-stability-with-type-as-argument-and-permutedims/108335 "2024-01-04T12:08:48Z")

</div>

I am experiencing poor performance (maybe type unstable?) when having a type as an optional argument to a function, and using permutedims inside. I’m having difficulty understanding why the first function is much slower …

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=29)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=31)
