# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=68

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 69

---

## [Speeding up contraction mapping](https://discourse.julialang.org/t/speeding-up-contraction-mapping/79134)

<div class="topic-metadata">

**Author:** [@rubaiyat](https://discourse.julialang.org/u/rubaiyat)\
**Replies:** 7\
**Last updated:** [April 8, 2022, 12:27am UTC](https://discourse.julialang.org/t/speeding-up-contraction-mapping/79134 "2022-04-08T00:27:21Z")

</div>

Hello everyone, I’m implementing the demand estimation for Berry-Levinsohn-Pakes (1995) for my research, and it would be great to get some help on this. While my code works, it takes a while, and so I was looking to spe…

---

## [Current OpenBLAS Versions (January 2022) do not support Intel gen 11 performantly?](https://discourse.julialang.org/t/current-openblas-versions-january-2022-do-not-support-intel-gen-11-performantly/75104)

<div class="topic-metadata">

**Author:** [@photor](https://discourse.julialang.org/u/photor)\
**Replies:** 50\
**Last updated:** [April 7, 2022, 2:35pm UTC](https://discourse.julialang.org/t/current-openblas-versions-january-2022-do-not-support-intel-gen-11-performantly/75104 "2022-04-07T14:35:27Z")

</div>

julia\> BLAS.get\_config() LinearAlgebra.BLAS.LBTConfig Libraries: └ \[ILP64\] libopenblas64\_.dll julia\> A=randn(10000,10000); julia\> @elapsed A\*A 31.8050192 julia\> @elapsed A\*A 32.776766 julia\> using MKL julia\> @elapse…

---

## [Reducing getindex allocation \[edit: when using an array as index\]](https://discourse.julialang.org/t/reducing-getindex-allocation-edit-when-using-an-array-as-index/79116)

<div class="topic-metadata">

**Author:** [@user\_231578](https://discourse.julialang.org/u/user_231578)\
**Replies:** 17\
**Last updated:** [April 7, 2022, 7:54am UTC](https://discourse.julialang.org/t/reducing-getindex-allocation-edit-when-using-an-array-as-index/79116 "2022-04-07T07:54:33Z")

</div>

I need to access a matrix millions of times, and I noticed that most of the allocations of my code come exactly from this. My code is similar to the following example const A = rand(10,10) function get\_number(idx) @inb…

---

## [Understanding @time memory allocations](https://discourse.julialang.org/t/understanding-time-memory-allocations/79075)

<div class="topic-metadata">

**Author:** [@qt\_codes](https://discourse.julialang.org/u/qt_codes)\
**Replies:** 7\
**Last updated:** [April 6, 2022, 5:44pm UTC](https://discourse.julialang.org/t/understanding-time-memory-allocations/79075 "2022-04-06T17:44:35Z")

</div>

I am timing a function in my code am confused as to why my memory allocations while compiling go up when I make the input smaller. The function is called twice, so I can look at pre and post compiling. @time begin m…

---

## [Efficiently interpreting byte-packed buffer](https://discourse.julialang.org/t/efficiently-interpreting-byte-packed-buffer/78975)

<div class="topic-metadata">

**Author:** [@nhardy](https://discourse.julialang.org/u/nhardy)\
**Replies:** 31\
**Last updated:** [April 6, 2022, 4:41pm UTC](https://discourse.julialang.org/t/efficiently-interpreting-byte-packed-buffer/78975 "2022-04-06T16:41:06Z")

</div>

tl;dr Is there a built-in way to interpret (or instantiate from) a byte-packed buffer to a Julia struct? If not, is there an efficient, non-allocating method to instantiate a new struct from the buffer data? I am repl…

---

## [Using Intel LLVM compiler with Julia](https://discourse.julialang.org/t/using-intel-llvm-compiler-with-julia/79065)

<div class="topic-metadata">

**Author:** [@duan](https://discourse.julialang.org/u/duan)\
**Replies:** 6\
**Last updated:** [April 5, 2022, 6:51pm UTC](https://discourse.julialang.org/t/using-intel-llvm-compiler-with-julia/79065 "2022-04-05T18:51:56Z")

</div>

Since Intel C/C++ compiler has adopted LLVM, is there a way to tell Julia to use Intel compiler?

---

## [@spawn at fails when called inside a function in a module](https://discourse.julialang.org/t/spawn-at-fails-when-called-inside-a-function-in-a-module/79001)

<div class="topic-metadata">

**Author:** [@Sijun](https://discourse.julialang.org/u/Sijun)\
**Replies:** 4\
**Last updated:** [April 5, 2022, 12:10am UTC](https://discourse.julialang.org/t/spawn-at-fails-when-called-inside-a-function-in-a-module/79001 "2022-04-05T00:10:39Z")

</div>

I have @spawnat expression run inside a function in a module and get the following error: module Test using Distributed function test() t = @spawnat :2 1==1 @show fetch(t) end end julia\> Test.test() ERROR: On…

---

## [Performance difference among three CUDA kernels with same results](https://discourse.julialang.org/t/performance-difference-among-three-cuda-kernels-with-same-results/78923)

<div class="topic-metadata">

**Author:** [@Chiil](https://discourse.julialang.org/u/Chiil)\
**Replies:** 8\
**Last updated:** [April 3, 2022, 3:23pm UTC](https://discourse.julialang.org/t/performance-difference-among-three-cuda-kernels-with-same-results/78923 "2022-04-03T15:23:41Z")

</div>

I wrote a small test program to learn a bit more of CUDA.jl. I find it hard to understand why the test3! function is so much slower than test1! and test2!, especially because this is the way one would write a CUDA kernel…

---

## [Is it necessary for sleep to allocate?](https://discourse.julialang.org/t/is-it-necessary-for-sleep-to-allocate/77795)

<div class="topic-metadata">

**Author:** [@taotree](https://discourse.julialang.org/u/taotree)\
**Replies:** 5\
**Last updated:** [April 2, 2022, 8:39pm UTC](https://discourse.julialang.org/t/is-it-necessary-for-sleep-to-allocate/77795 "2022-04-02T20:39:30Z")

</div>

I was optimizing a complex block of multithreaded code and discovered that sleep was allocating. julia\> @btime sleep(0) 1.720 μs (5 allocations: 144 bytes) I was pretty surprised, so checked the source which is basic…

---

## [Maintaining a fixed size "top N" values list](https://discourse.julialang.org/t/maintaining-a-fixed-size-top-n-values-list/78868)

<div class="topic-metadata">

**Author:** [@taotree](https://discourse.julialang.org/u/taotree)\
**Replies:** 14\
**Last updated:** [April 2, 2022, 3:55pm UTC](https://discourse.julialang.org/t/maintaining-a-fixed-size-top-n-values-list/78868 "2022-04-02T15:55:50Z")

</div>

My high level need: loop through lots of entries and keep track of the top nth value so I can continue filtering out anything lower than the nth highest value. My assumption is that I would need to store the n top value…

---

## [Function array performance](https://discourse.julialang.org/t/function-array-performance/78889)

<div class="topic-metadata">

**Author:** [@smith-isaac](https://discourse.julialang.org/u/smith-isaac)\
**Replies:** 5\
**Last updated:** [April 2, 2022, 2:05pm UTC](https://discourse.julialang.org/t/function-array-performance/78889 "2022-04-02T14:05:26Z")

</div>

I have an array of functions that each return a float, and so I am trying to assign the results of those functions to an array. So I had something like this results = similar(f\_arr, Float64) results .= \[f(a, b, c) for f…

---

## [Poll: speed vs accuracy for \`Float64^-3\`](https://discourse.julialang.org/t/poll-speed-vs-accuracy-for-float64-3/77619)

<div class="topic-metadata">

**Author:** [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Replies:** 55\
**Last updated:** [March 31, 2022, 4:01pm UTC](https://discourse.julialang.org/t/poll-speed-vs-accuracy-for-float64-3/77619 "2022-03-31T16:01:23Z")

</div>

In Julia 1.8, we are getting a pure Julia pow function, and as discussed in https://github.com/JuliaLang/julia/pull/44529, part of this involved fixing a bug in x^-3. Before 1.8, Julia’s Literal pow meant that 10^-3 = 0…

---

## [Memory allocation when evaluating a function inside struct](https://discourse.julialang.org/t/memory-allocation-when-evaluating-a-function-inside-struct/78784)

<div class="topic-metadata">

**Author:** [@Kishore\_Nori](https://discourse.julialang.org/u/Kishore_Nori)\
**Replies:** 6\
**Last updated:** [March 31, 2022, 2:19pm UTC](https://discourse.julialang.org/t/memory-allocation-when-evaluating-a-function-inside-struct/78784 "2022-03-31T14:19:51Z")

</div>

I am trying to understand the following memory allocation behaviour: struct justfunc f::Function end func = justfunc(sin) @allocated func.f(0.1) # 32 @allocated sin(0.1) # 0 Having an other variant where the type i…

---

## [Memory allocation in iterator](https://discourse.julialang.org/t/memory-allocation-in-iterator/78786)

<div class="topic-metadata">

**Author:** [@Tqft](https://discourse.julialang.org/u/Tqft)\
**Replies:** 9\
**Last updated:** [March 31, 2022, 9:40am UTC](https://discourse.julialang.org/t/memory-allocation-in-iterator/78786 "2022-03-31T09:40:54Z")

</div>

There are already several discussions about memory allocations in iterators but I have a very simple question that bothers me whenever I implement an iterator. Take the example iterator from the manual: julia\> struct Sq…

---

## [Improving my mental model: Why is \`sizehint!()\` so much slower than \`zeros()\`?](https://discourse.julialang.org/t/improving-my-mental-model-why-is-sizehint-so-much-slower-than-zeros/78776)

<div class="topic-metadata">

**Author:** [@maxkapur](https://discourse.julialang.org/u/maxkapur)\
**Replies:** 9\
**Last updated:** [March 31, 2022, 9:03am UTC](https://discourse.julialang.org/t/improving-my-mental-model-why-is-sizehint-so-much-slower-than-zeros/78776 "2022-03-31T09:03:57Z")

</div>

I am asking this question in an effort to improve my understanding of computer science, not to complain about Julia’s performance :wink: I understand that, in principle, if we know that we are going to be filling an vec…

---

## [Super slow string performance](https://discourse.julialang.org/t/super-slow-string-performance/78156)

<div class="topic-metadata">

**Author:** [@peteole](https://discourse.julialang.org/u/peteole)\
**Replies:** 2\
**Last updated:** [March 31, 2022, 8:59am UTC](https://discourse.julialang.org/t/super-slow-string-performance/78156 "2022-03-31T08:59:46Z")

</div>

Hi, while solving a coding competition problem with julia I experienced extreme performance issues with the String type. This program: function deleteCount(target::String, input::String) if length(target) \> length…

---

## [Avoiding double lookup](https://discourse.julialang.org/t/avoiding-double-lookup/78636)

<div class="topic-metadata">

**Author:** [@jar1](https://discourse.julialang.org/u/jar1)\
**Replies:** 11\
**Last updated:** [March 31, 2022, 12:18am UTC](https://discourse.julialang.org/t/avoiding-double-lookup/78636 "2022-03-31T00:18:58Z")

</div>

Is it possible to cache the location so I don’t need to look it up twice (once in get and once in setindex!)? "Increment the count by 1, starting from 0." function bump!(d::Dict{K,Int}, k::K) where K v = get(d, k, 0…

---

## [Can the output of --track-allocation be trusted with multi-threading enabled?](https://discourse.julialang.org/t/can-the-output-of-track-allocation-be-trusted-with-multi-threading-enabled/78744)

<div class="topic-metadata">

**Author:** [@krcools](https://discourse.julialang.org/u/krcools)\
**Replies:** 9\
**Last updated:** [March 30, 2022, 6:07pm UTC](https://discourse.julialang.org/t/can-the-output-of-track-allocation-be-trusted-with-multi-threading-enabled/78744 "2022-03-30T18:07:43Z")

</div>

Below the output of --track-allocation. The first time when the code is ran with a single thread, the second time when the work is divided and the code is called from 8 threads. Is this output to be trusted? If so where …

---

## [When to parallelise, when not to?](https://discourse.julialang.org/t/when-to-parallelise-when-not-to/78722)

<div class="topic-metadata">

**Author:** [@ayushinav](https://discourse.julialang.org/u/ayushinav)\
**Replies:** 2\
**Last updated:** [March 30, 2022, 7:30am UTC](https://discourse.julialang.org/t/when-to-parallelise-when-not-to/78722 "2022-03-30T07:30:17Z")

</div>

Hey all, I’m stuck on understanding when to realize that the code will benefit from parallel processing, when not; when to take it on GPUs when not. I have a code that runs as fast as it can (that’s what I think, could…

---

## [Get NamedTuple element with String key. Why is this slow?](https://discourse.julialang.org/t/get-namedtuple-element-with-string-key-why-is-this-slow/78704)

<div class="topic-metadata">

**Author:** [@jocklawrie](https://discourse.julialang.org/u/jocklawrie)\
**Replies:** 3\
**Last updated:** [March 30, 2022, 5:26am UTC](https://discourse.julialang.org/t/get-namedtuple-element-with-string-key-why-is-this-slow/78704 "2022-03-30T05:26:47Z")

</div>

Hi there, I’m attempting to get a value from a NamedTuple using a String key, ideally using x\[k\] syntax. In the example below, why is x\["a"\] so slow compared to x\[Symbol("a")\]? using BenchmarkTools import Base: getinde…

---

## [In Nearest Neighbours are trees other than BruteTree worth it for sets with between n=2^3 and n=2^20ish for KNN in 2D cartesian space](https://discourse.julialang.org/t/in-nearest-neighbours-are-trees-other-than-brutetree-worth-it-for-sets-with-between-n-2-3-and-n-2-20ish-for-knn-in-2d-cartesian-space/78561)

<div class="topic-metadata">

**Author:** [@AlanAylmer](https://discourse.julialang.org/u/AlanAylmer)\
**Replies:** 2\
**Last updated:** [March 28, 2022, 6:15pm UTC](https://discourse.julialang.org/t/in-nearest-neighbours-are-trees-other-than-brutetree-worth-it-for-sets-with-between-n-2-3-and-n-2-20ish-for-knn-in-2d-cartesian-space/78561 "2022-03-28T18:15:41Z")

</div>

I know the colossal @kristoffer.carlsson has been busy recently but I wonder if anyone else knows the A to my Q… I am running a teaching group on grouping waveforms. The waveform side of it is okay but I want to unders…

---

## [CUDA kernel configuration](https://discourse.julialang.org/t/cuda-kernel-configuration/78417)

<div class="topic-metadata">

**Author:** [@Marcell\_Havlik](https://discourse.julialang.org/u/Marcell_Havlik)\
**Replies:** 3\
**Last updated:** [March 28, 2022, 7:54am UTC](https://discourse.julialang.org/t/cuda-kernel-configuration/78417 "2022-03-28T07:54:41Z")

</div>

Hey Julianners, I don’t know what do I do wrong, but I use a configurator that calculates the ideal kernel config that was used here: https://discourse.julialang.org/t/the-most-general-way-to-estimate-the-optimal-argume…

---

## [Type instability of nested recursive function](https://discourse.julialang.org/t/type-instability-of-nested-recursive-function/78392)

<div class="topic-metadata">

**Author:** [@ma-chengyuan](https://discourse.julialang.org/u/ma-chengyuan)\
**Replies:** 4\
**Last updated:** [March 28, 2022, 3:38am UTC](https://discourse.julialang.org/t/type-instability-of-nested-recursive-function/78392 "2022-03-28T03:38:10Z")

</div>

I was writing some code where I need an inner function to be recursive, something like (highly simplified) function f1(x::Vector{Int64}) function g(i, val) i \<= length(x) || return g(i + 1, val) …

---

## [\# of threads in Optimal Transport Package](https://discourse.julialang.org/t/of-threads-in-optimal-transport-package/78560)

<div class="topic-metadata">

**Author:** [@kadir-gunel](https://discourse.julialang.org/u/kadir-gunel)\
**Replies:** 4\
**Last updated:** [March 27, 2022, 6:38pm UTC](https://discourse.julialang.org/t/of-threads-in-optimal-transport-package/78560 "2022-03-27T18:38:25Z")

</div>

Hello, I tried to increase the # of threads in OT.jl by initializing julia with flags such as -p 16 or -t 32 but no success. I can only use 8 threads by default. Is there any way that we can increase the # of threads ? …

---

## [Will Julia be more efficient than PyTorch](https://discourse.julialang.org/t/will-julia-be-more-efficient-than-pytorch/78557)

<div class="topic-metadata">

**Author:** [@mindaslab](https://discourse.julialang.org/u/mindaslab)\
**Replies:** 1\
**Last updated:** [March 27, 2022, 5:29pm UTC](https://discourse.julialang.org/t/will-julia-be-more-efficient-than-pytorch/78557 "2022-03-27T17:29:19Z")

</div>

Hello All, I am very much interested in deep learning, but I realize that I cannot buy expensive hardware that industries use to train things like GPT-NeoX and so on.I wonder if a neural network is built and trained on…

---

## [Extremely slow n-th root of float](https://discourse.julialang.org/t/extremely-slow-n-th-root-of-float/78534)

<div class="topic-metadata">

**Author:** [@goerz](https://discourse.julialang.org/u/goerz)\
**Replies:** 11\
**Last updated:** [March 27, 2022, 5:23am UTC](https://discourse.julialang.org/t/extremely-slow-n-th-root-of-float/78534 "2022-03-27T05:23:18Z")

</div>

I’m benchmarking an algorithm translated from Fortran, doing something similar to Expokit.expmv – evaluate the result of exponentiation a matrix and applying it to a vector, by expanding the exp into a polynomial series. …

---

## [Julia 3 times slower than Fortran reading integer data from ASCII file](https://discourse.julialang.org/t/julia-3-times-slower-than-fortran-reading-integer-data-from-ascii-file/78516)

<div class="topic-metadata">

**Author:** [@jman87](https://discourse.julialang.org/u/jman87)\
**Replies:** 14\
**Last updated:** [March 26, 2022, 7:20pm UTC](https://discourse.julialang.org/t/julia-3-times-slower-than-fortran-reading-integer-data-from-ascii-file/78516 "2022-03-26T19:20:13Z")

</div>

Hey all, I’ve been dabbling with Julia since the 0. days. Up to point I haven’t had any real performance issues, but am now needing to read in some fairly large size files (think finite element / mesh data with millions …

---

## [Type-instability because of @threads boxing variables](https://discourse.julialang.org/t/type-instability-because-of-threads-boxing-variables/78395)

<div class="topic-metadata">

**Author:** [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Replies:** 23\
**Last updated:** [March 25, 2022, 11:38pm UTC](https://discourse.julialang.org/t/type-instability-because-of-threads-boxing-variables/78395 "2022-03-25T23:38:29Z")

</div>

See Type-instability because of @threads boxing variables - #12 by lmiq for a more precise description of the problem. Why are allocations occurring (or being reported, at least, by @allocated) in this example, sort of…

---

## [Combining lots of DataFrames, best approach?](https://discourse.julialang.org/t/combining-lots-of-dataframes-best-approach/78466)

<div class="topic-metadata">

**Author:** [@tbeason](https://discourse.julialang.org/u/tbeason)\
**Replies:** 10\
**Last updated:** [March 25, 2022, 6:05pm UTC](https://discourse.julialang.org/t/combining-lots-of-dataframes-best-approach/78466 "2022-03-25T18:05:38Z")

</div>

I am wondering if there are better ways to append lots of DataFrames together in these two scenarios. # dts is a vector of lots of dataframes (\>1000) # here is a MWE version of it dts = \[DataFrame(pid=fill(string(rand(U…

---

## [How to preallocate a triangular matrix and overwrite it with cholesky!](https://discourse.julialang.org/t/how-to-preallocate-a-triangular-matrix-and-overwrite-it-with-cholesky/78463)

<div class="topic-metadata">

**Author:** [@LuZhangstat](https://discourse.julialang.org/u/LuZhangstat)\
**Replies:** 9\
**Last updated:** [March 25, 2022, 3:37pm UTC](https://discourse.julialang.org/t/how-to-preallocate-a-triangular-matrix-and-overwrite-it-with-cholesky/78463 "2022-03-25T15:37:29Z")

</div>

Hi everyone, I am trying to improve my code’s efficiency. I locate a line that might be too expensive, but I failed to figure out how to improve it. In brief, I need to save the inverse of an N by N covariance matrix of…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=67)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=69)
