# Performance

**URL:** https://discourse.julialang.org/c/usage/perf/37.md?page=72

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

**Page:** 73

---

## [Multithreaded code on beefy computer runs just as fast as serial code on M1 Mac](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210)

<div class="topic-metadata">

**Author:** [@ash](https://discourse.julialang.org/u/ash)\
**Replies:** 41\
**Last updated:** [February 2, 2022, 1:10pm UTC](https://discourse.julialang.org/t/multithreaded-code-on-beefy-computer-runs-just-as-fast-as-serial-code-on-m1-mac/75210 "2022-02-02T13:10:16Z")

</div>

I have 2 machines: M1 Mac Mini, with horrible thermals and a peak of 3.2 GHz and 8 GB of RAM (absolutely love it though) A beefy linux (Ubuntu 21.10) workstation with 64 GB of DDR5 RAM and a 12th-Gen 4.9 GHz processor …

---

## [Rust / Julia Comparison Post](https://discourse.julialang.org/t/rust-julia-comparison-post/75403)

<div class="topic-metadata">

**Author:** [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Replies:** 56\
**Last updated:** [February 1, 2022, 11:38pm UTC](https://discourse.julialang.org/t/rust-julia-comparison-post/75403 "2022-02-01T23:38:10Z")

</div>

This post is on Hacker News right now – it seems worth digging in to understand why Rust does so much better here given that the code seem like a fairly naive port without much optimization effort: Julia Rust

---

## [Performance of custom \`Vec\` type versus \`SVector{3, Float64}\`](https://discourse.julialang.org/t/performance-of-custom-vec-type-versus-svector-3-float64/75594)

<div class="topic-metadata">

**Author:** [@CameronBieganek](https://discourse.julialang.org/u/CameronBieganek)\
**Replies:** 4\
**Last updated:** [February 1, 2022, 7:47pm UTC](https://discourse.julialang.org/t/performance-of-custom-vec-type-versus-svector-3-float64/75594 "2022-02-01T19:47:05Z")

</div>

I’m having trouble replicating the performance of SVector{3, Float64} with my own custom Vec type. Here’s the code: using LinearAlgebra using StaticArrays using BenchmarkTools struct Vec x::Float64 y::Float64 …

---

## [How to control threads in combination of LoopVectorization and @spawn](https://discourse.julialang.org/t/how-to-control-threads-in-combination-of-loopvectorization-and-spawn/75466)

<div class="topic-metadata">

**Author:** [@Chiil](https://discourse.julialang.org/u/Chiil)\
**Replies:** 6\
**Last updated:** [February 1, 2022, 7:26pm UTC](https://discourse.julialang.org/t/how-to-control-threads-in-combination-of-loopvectorization-and-spawn/75466 "2022-02-01T19:26:56Z")

</div>

I am having a simple test program to learn how to divide work over threads. If I launch with one thread, the output is: ──────────────────────────────────────────────────────────────────── Ti…

---

## [Parallelizing a Nested Loop Concurrently](https://discourse.julialang.org/t/parallelizing-a-nested-loop-concurrently/75402)

<div class="topic-metadata">

**Author:** [@jmcastro2109](https://discourse.julialang.org/u/jmcastro2109)\
**Replies:** 17\
**Last updated:** [February 1, 2022, 5:23pm UTC](https://discourse.julialang.org/t/parallelizing-a-nested-loop-concurrently/75402 "2022-02-01T17:23:03Z")

</div>

Hi! I am trying to code a problem in which I have two loops. An inner loop that a computes a sample average for a given agent using many many simulations of a particular function (which is not trivial, but each individua…

---

## [Lookup in Dict{Int,Float64} allocates](https://discourse.julialang.org/t/lookup-in-dict-int-float64-allocates/75495)

<div class="topic-metadata">

**Author:** [@johnomotani](https://discourse.julialang.org/u/johnomotani)\
**Replies:** 2\
**Last updated:** [January 31, 2022, 12:27pm UTC](https://discourse.julialang.org/t/lookup-in-dict-int-float64-allocates/75495 "2022-01-31T12:27:12Z")

</div>

I was surprised that the following code reports an allocation when looking up a Float64 value from a Dict{Int64,Float64} using BenchmarkTools n = 1\_000\_000 d = Dict(i=\>rand() for i in 1:n); @btime d\[42\] # 19.346 ns (1 …

---

## [Regarding the multithreaded performance of OpenBLAS](https://discourse.julialang.org/t/regarding-the-multithreaded-performance-of-openblas/75450)

<div class="topic-metadata">

**Author:** [@ash](https://discourse.julialang.org/u/ash)\
**Replies:** 7\
**Last updated:** [January 31, 2022, 6:33am UTC](https://discourse.julialang.org/t/regarding-the-multithreaded-performance-of-openblas/75450 "2022-01-31T06:33:25Z")

</div>

I have been benchmarking some code and coming across all sorts of questions (see my other thread for in-depth details, I will post a progress update soon). However, I have an OpenBLAS question that I believe merits its …

---

## [GLM - inconsistency with R on ISLR dataset?](https://discourse.julialang.org/t/glm-inconsistency-with-r-on-islr-dataset/75427)

<div class="topic-metadata">

**Author:** [@compleat](https://discourse.julialang.org/u/compleat)\
**Replies:** 2\
**Last updated:** [January 30, 2022, 5:31pm UTC](https://discourse.julialang.org/t/glm-inconsistency-with-r-on-islr-dataset/75427 "2022-01-30T17:31:11Z")

</div>

I am teaching a class using the (well-known) book Intro to Statistical Learning in R (ISLR). For the Lab example in Chapter 3, there is an example using polynomial regression in GLM. According to documentation, the wa…

---

## [More threads, slower code, even if not spawning them](https://discourse.julialang.org/t/more-threads-slower-code-even-if-not-spawning-them/75242)

<div class="topic-metadata">

**Author:** [@lmiq](https://discourse.julialang.org/u/lmiq)\
**Replies:** 19\
**Last updated:** [January 29, 2022, 9:16pm UTC](https://discourse.julialang.org/t/more-threads-slower-code-even-if-not-spawning-them/75242 "2022-01-29T21:16:37Z")

</div>

Maybe I don’t understand how @spawn works, but here is what I am observing: When executing a parallel task with it, I observe that the performance is significantly worse if I have more threads available (i. e. starting j…

---

## [Counting all same path in a JSON](https://discourse.julialang.org/t/counting-all-same-path-in-a-json/75433)

<div class="topic-metadata">

**Author:** [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Replies:** 0\
**Last updated:** [January 29, 2022, 8:50pm UTC](https://discourse.julialang.org/t/counting-all-same-path-in-a-json/75433 "2022-01-29T20:50:07Z")

</div>

Hi All, I am curios if someone was solving a following problem. In JsonGrinder.jl, I need to store leafs of JSONs of the same type (same path) to the same continuous array, such that when I pass it to Neural Network, B…

---

## [Conversion of CxxWrap.StdLib.StdVectorAllocated{UInt8} to Julian Vector{UInt8} faster](https://discourse.julialang.org/t/conversion-of-cxxwrap-stdlib-stdvectorallocated-uint8-to-julian-vector-uint8-faster/75351)

<div class="topic-metadata">

**Author:** [@Ashwani\_Rathee](https://discourse.julialang.org/u/Ashwani_Rathee)\
**Replies:** 8\
**Last updated:** [January 29, 2022, 6:21am UTC](https://discourse.julialang.org/t/conversion-of-cxxwrap-stdlib-stdvectorallocated-uint8-to-julian-vector-uint8-faster/75351 "2022-01-29T06:21:02Z")

</div>

Hey guys, issue I am facing is the conversion of the arrays and the efficiency of the conversion julia\> res = screenshot\_rgb() # image of size 1920\*1080 with 4 bytes per pixel returned 8294400-element CxxWrap.StdLib.Std…

---

## [Evaluating @view macro before @tturbo](https://discourse.julialang.org/t/evaluating-view-macro-before-tturbo/75320)

<div class="topic-metadata">

**Author:** [@Chiil](https://discourse.julialang.org/u/Chiil)\
**Replies:** 1\
**Last updated:** [January 27, 2022, 11:54pm UTC](https://discourse.julialang.org/t/evaluating-view-macro-before-tturbo/75320 "2022-01-27T23:54:04Z")

</div>

I am revising a piece of code where I am removing unnecessary copies by creating views. At places where I have @tturbo decorators, I get errors because LoopVectorization does not like the @view. Is it possible to make th…

---

## [Could this be faster using some decomposition?](https://discourse.julialang.org/t/could-this-be-faster-using-some-decomposition/75323)

<div class="topic-metadata">

**Author:** [@lrnv](https://discourse.julialang.org/u/lrnv)\
**Replies:** 4\
**Last updated:** [January 27, 2022, 10:02pm UTC](https://discourse.julialang.org/t/could-this-be-faster-using-some-decomposition/75323 "2022-01-27T22:02:00Z")

</div>

Hi, This is more maths than Julia optimisation, but I do not recall if there is a trick in the linear algebra book to make this more efficiently. I have a covariance matrix M, of size (2000,2000), and i want to normali…

---

## [Evaluating the Hessian of the Lagrangian efficiently](https://discourse.julialang.org/t/evaluating-the-hessian-of-the-lagrangian-efficiently/54659)

<div class="topic-metadata">

**Author:** [@ferrolho](https://discourse.julialang.org/u/ferrolho)\
**Replies:** 14\
**Last updated:** [January 27, 2022, 9:49am UTC](https://discourse.julialang.org/t/evaluating-the-hessian-of-the-lagrangian-efficiently/54659 "2022-01-27T09:49:31Z")

</div>

In the context of numerical optimisation, it is good to provide exact Hessians of the problem constraints. I have a working version thanks to Julia’s AD capabilities, but I think my current approach is not optimised for …

---

## [Allocation with Union Nothing](https://discourse.julialang.org/t/allocation-with-union-nothing/75163)

<div class="topic-metadata">

**Author:** [@RomainPct](https://discourse.julialang.org/u/RomainPct)\
**Replies:** 2\
**Last updated:** [January 26, 2022, 1:12am UTC](https://discourse.julialang.org/t/allocation-with-union-nothing/75163 "2022-01-26T01:12:17Z")

</div>

Hello, I discovered this behavior that created some weird allocations in my code. I feel like I get it now, but I would be interested in some relevant comments/explanations. f\_allocates(x) = isempty(x) ? nothing : (0, …

---

## [Improving performance of dynamic programming problem of households in a labor market model](https://discourse.julialang.org/t/improving-performance-of-dynamic-programming-problem-of-households-in-a-labor-market-model/74892)

<div class="topic-metadata">

**Author:** [@Mario\_Silva](https://discourse.julialang.org/u/Mario_Silva)\
**Replies:** 25\
**Last updated:** [January 26, 2022, 12:13am UTC](https://discourse.julialang.org/t/improving-performance-of-dynamic-programming-problem-of-households-in-a-labor-market-model/74892 "2022-01-26T00:13:04Z")

</div>

Greetings: I am working on solving a heterogeneous-agent economics model in Julia. The most demanding part is solving the household problem using dynamic programming. Households can be either employed or unemployed. The…

---

## [Help optimizing bad multithreading code](https://discourse.julialang.org/t/help-optimizing-bad-multithreading-code/75099)

<div class="topic-metadata">

**Author:** [@isaac.rubberducky](https://discourse.julialang.org/u/isaac.rubberducky)\
**Replies:** 2\
**Last updated:** [January 25, 2022, 2:41pm UTC](https://discourse.julialang.org/t/help-optimizing-bad-multithreading-code/75099 "2022-01-25T14:41:00Z")

</div>

function cleanExample(rootNode::Node) queue = Vector{Node}() push!(queue, rootNode) bestKnownSolution = 100 queueLock = ReentrantLock() solutionLock = ReentrantLock() while !isempty(queue) …

---

## [Help with setting up Distributed computing](https://discourse.julialang.org/t/help-with-setting-up-distributed-computing/75081)

<div class="topic-metadata">

**Author:** [@isaac.rubberducky](https://discourse.julialang.org/u/isaac.rubberducky)\
**Replies:** 5\
**Last updated:** [January 25, 2022, 2:39pm UTC](https://discourse.julialang.org/t/help-with-setting-up-distributed-computing/75081 "2022-01-25T14:39:04Z")

</div>

queue = Vector{Node}() push!(queue, rootNode) bestKnownSolution = 100 while !isempty(queue) nextNode = pop!(queue) newNode1, newNode2, bestFoundSolution = processNode!(nextNode) if !isnothing(newNode1) …

---

## [Best Julian way to copy many files?](https://discourse.julialang.org/t/best-julian-way-to-copy-many-files/75142)

<div class="topic-metadata">

**Author:** [@chowerth](https://discourse.julialang.org/u/chowerth)\
**Replies:** 2\
**Last updated:** [January 24, 2022, 7:24pm UTC](https://discourse.julialang.org/t/best-julian-way-to-copy-many-files/75142 "2022-01-24T19:24:00Z")

</div>

Hello. I’m wondering what is the best Julian way to copy files. My particular use case is I have a large number of files (of varying size) and I’m automating their data cleaning and backup process. Without going into gr…

---

## [Speeding up maximization](https://discourse.julialang.org/t/speeding-up-maximization/74298)

<div class="topic-metadata">

**Author:** [@juliauser](https://discourse.julialang.org/u/juliauser)\
**Replies:** 4\
**Last updated:** [January 24, 2022, 4:12am UTC](https://discourse.julialang.org/t/speeding-up-maximization/74298 "2022-01-24T04:12:15Z")

</div>

I am trying to speed up this code. The maximization is very time consuming. Perhaps I am also not doing the allocations efficiently. I have a core with 64 threads, but if I use multithread I incur in an error (the result…

---

## [PyCall stdout buffer processing not doing what I expected](https://discourse.julialang.org/t/pycall-stdout-buffer-processing-not-doing-what-i-expected/75086)

<div class="topic-metadata">

**Author:** [@anon69491625](https://discourse.julialang.org/u/anon69491625)\
**Replies:** 10\
**Last updated:** [January 23, 2022, 9:37pm UTC](https://discourse.julialang.org/t/pycall-stdout-buffer-processing-not-doing-what-i-expected/75086 "2022-01-23T21:37:57Z")

</div>

Hi all I am using PyCall to run a simple python script that prints out the ascii table. #test\_stdout\_pycall.py digit = 0 while digit\<=255: print("%d = %c" %(digit, chr(digit))) digit += 1 print("done") I w…

---

## [MacOS ARM64 no faster than emulated x86?](https://discourse.julialang.org/t/macos-arm64-no-faster-than-emulated-x86/72481)

<div class="topic-metadata">

**Author:** [@dodoplus](https://discourse.julialang.org/u/dodoplus)\
**Replies:** 17\
**Last updated:** [January 22, 2022, 8:37pm UTC](https://discourse.julialang.org/t/macos-arm64-no-faster-than-emulated-x86/72481 "2022-01-22T20:37:45Z")

</div>

I was curious to test the performance of the newly released Julia 1.7 running natively on my M1 laptop. Surprisingly, there doesn’t seem to be any speedup: julia\> include("b.jl") sortperf (generic function with 1 metho…

---

## [Repeated multithreaded region](https://discourse.julialang.org/t/repeated-multithreaded-region/75037)

<div class="topic-metadata">

**Author:** [@bjack205](https://discourse.julialang.org/u/bjack205)\
**Replies:** 2\
**Last updated:** [January 22, 2022, 4:13pm UTC](https://discourse.julialang.org/t/repeated-multithreaded-region/75037 "2022-01-22T16:13:28Z")

</div>

I’m wondering what the best way to repeatedly call a parallelized piece of code in a serial loop? For example, in smooth nonlinear optimization, we often need to calculate an expansion of objective and constraints about …

---

## [Reducing the memory allocation in ODEs](https://discourse.julialang.org/t/reducing-the-memory-allocation-in-odes/74402)

<div class="topic-metadata">

**Author:** [@xspeng](https://discourse.julialang.org/u/xspeng)\
**Replies:** 11\
**Last updated:** [January 22, 2022, 2:36am UTC](https://discourse.julialang.org/t/reducing-the-memory-allocation-in-odes/74402 "2022-01-22T02:36:53Z")

</div>

Just coming from Matlab, to pursue faster speed. Here is my problem and a simple case Even run for the second time(after compilation), there are still thousands of allocation, any suggestions are welcome function Eq(dy…

---

## [Matrix multiplication is slower when multithreading in Julia](https://discourse.julialang.org/t/matrix-multiplication-is-slower-when-multithreading-in-julia/56227)

<div class="topic-metadata">

**Author:** [@pablof300](https://discourse.julialang.org/u/pablof300)\
**Replies:** 13\
**Last updated:** [January 21, 2022, 10:26pm UTC](https://discourse.julialang.org/t/matrix-multiplication-is-slower-when-multithreading-in-julia/56227 "2022-01-21T22:26:46Z")

</div>

I am working with big matrices (size of 30k rows and ~100 columns). I am doing some matrix multiplication and the process would take around 20 seconds. This is my code: @time begin result = -1 data =…

---

## [Adding a sysimage to a julia script - Package Complier](https://discourse.julialang.org/t/adding-a-sysimage-to-a-julia-script-package-complier/74994)

<div class="topic-metadata">

**Author:** [@maulik.bhatt](https://discourse.julialang.org/u/maulik.bhatt)\
**Replies:** 0\
**Last updated:** [January 21, 2022, 4:56pm UTC](https://discourse.julialang.org/t/adding-a-sysimage-to-a-julia-script-package-complier/74994 "2022-01-21T16:56:41Z")

</div>

I have a Julia script which uses several packages and runs several functions from them. This Julia script takes a lot of time every time I run it because every package is precompiled in julia every time I run that script…

---

## [Weighted histogram ~2x as slow in Julia vs. Python](https://discourse.julialang.org/t/weighted-histogram-2x-as-slow-in-julia-vs-python/74953)

<div class="topic-metadata">

**Author:** [@kirklong](https://discourse.julialang.org/u/kirklong)\
**Replies:** 13\
**Last updated:** [January 20, 2022, 9:36pm UTC](https://discourse.julialang.org/t/weighted-histogram-2x-as-slow-in-julia-vs-python/74953 "2022-01-20T21:36:35Z")

</div>

I’m working on a project where I need to compute the weighted histogram of something many times. Basically I need the sum of each of the bins of the histogram, and I’ve been using PyCall to call the SciPy “binned statist…

---

## [Package set-up not propagating to workers with Distributed](https://discourse.julialang.org/t/package-set-up-not-propagating-to-workers-with-distributed/74952)

<div class="topic-metadata">

**Author:** [@danielsoutar](https://discourse.julialang.org/u/danielsoutar)\
**Replies:** 0\
**Last updated:** [January 20, 2022, 6:09pm UTC](https://discourse.julialang.org/t/package-set-up-not-propagating-to-workers-with-distributed/74952 "2022-01-20T18:09:03Z")

</div>

Info: $ julia --version julia version 1.6.0 $ lscpu ~/root/MyPackage$ lscpu Architecture: x86\_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian Address …

---

## [Multithreading for nested loops](https://discourse.julialang.org/t/multithreading-for-nested-loops/36002)

<div class="topic-metadata">

**Author:** [@henry2004y](https://discourse.julialang.org/u/henry2004y)\
**Replies:** 42\
**Last updated:** [January 20, 2022, 6:00pm UTC](https://discourse.julialang.org/t/multithreading-for-nested-loops/36002 "2022-01-20T18:00:37Z")

</div>

Hi, I am wondering how to apply multi-threading to nested for loops. For the function nestedloop below, I guess the threads will only be launched for the outermost k loop for OMP\_JULIA\_THREADS times: function nestedlo…

---

## [Pmap extremely slow when function returns large object](https://discourse.julialang.org/t/pmap-extremely-slow-when-function-returns-large-object/74937)

<div class="topic-metadata">

**Author:** [@user\_231578](https://discourse.julialang.org/u/user_231578)\
**Replies:** 4\
**Last updated:** [January 20, 2022, 2:20pm UTC](https://discourse.julialang.org/t/pmap-extremely-slow-when-function-returns-large-object/74937 "2022-01-20T14:20:18Z")

</div>

I would like to parallelize a function but returning large objects makes pmap extremely slow. using Distributed using SharedArrays addprocs(length(Sys.cpu\_info())-1) @everywhere using Random @everywhere function foo(i…

[Previous page](https://discourse.julialang.org/c/usage/perf/37.md?page=71)

[Next page](https://discourse.julialang.org/c/usage/perf/37.md?page=73)
