# Tracking memory usage in unit tests (a lot worse in Julia 0.7 than 0.6)

**URL:** <https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444>\
**Category:** General Usage\
**Created:** [September 2, 2018, 6:55pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444 "2018-09-02T18:55:17Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [September 2, 2018, 6:55pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/1 "2018-09-02T18:55:17Z")

</div>

Upgrading our package from Julia 0.6 to 0.7, we’ve noticed that unit tests use a lot more memory. Trivial tests that should in theory allocate a few KB can now allocate hundreds of MB. A minimal example:

```julia
function test()
    A = [x for x=0.0:2]
end

test()

```

Running it:

```julia
julia> @time include("test.jl")
  0.236493 seconds (730.15 k allocations: 38.490 MiB, 2.46% gc time)

```

I understand that the usual answer to this type of question is to time the code twice, since the first run includes JIT, etc. And indeed, if I run `@time test()` twice, the second run gives a more reasonable `0.000004 seconds (5 allocations: 272 bytes)`. But it doesn’t seem very practical when running a large number of unit tests to have to run each test twice (some of our tests take quite a bit of time). Is there a recommended way of doing this? Our goals are 1) somewhat accurate performance measurements so that we can detect regressions, and 2) minimize total time spent running test suite.

In particular, we can’t really tell if we’ve regressed since Julia 0.6, since first invocation of code seems to take a lot longer and use much more memory in Julia 0.7 than in 0.6. Same code above with Julia 0.6.4:

```julia
julia> @time include("test.jl")
  0.111263 seconds (60.98 k allocations: 3.406 MiB)

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 2, 2018, 7:14pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/2 "2018-09-02T19:14:55Z")

</div>

Are you sure you eliminated all deprecation warnings? Sometimes they may not show explicitly. Run `julia` with the option `--depwarn=yes` to get errors instead.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [September 3, 2018, 7:21am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/3 "2018-09-03T07:21:14Z")

</div>

Yes, all deprecation warnings are removed. But it’s easy to reproduce the problem I’m referring to with a minimal unit test not related to our package:

```julia
using Test

function array()
    [x for x=0.0:2.0]
end

@test all(array() .< 10)

```

Running this test:

```julia
julia> @time include("test.jl")
  0.681974 seconds (2.28 M allocations: 114.947 MiB, 3.20% gc time)
Test Passed

```

With this kind of numbers, it becomes almost pointless to measure performance and memory usage, since I guess we’re measuring JIT compile time and not run time of the code tested.

So how do package developers solve this to track performance in their tests? The best idea I can think of is to alter each test to run twice, and do the timing on the second run only, and from within the test instead of outside, i.e.:

```julia
using Test

function array()
    [x for x=0.0:2.0]
end

@test all(array() .< 10)
@time all(array() .< 10)

```

Which results in more accurate numbers:

```julia
julia> include("test.jl")
  0.000015 seconds (10 allocations: 4.625 KiB)

```

But this means that we have to run our entire test suite twice, and it’s already slow enough as it is.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 3, 2018, 7:47am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/4 "2018-09-03T07:47:27Z")

</div>

> [@bennedich](#):
>
> With this kind of numbers, it becomes almost pointless to measure performance and memory usage, since I guess we’re measuring JIT compile time and not run time of the code tested.
> 
> So how do package developers solve this to track performance in their tests?

> **[GitHub - JuliaCI/BenchmarkTools.jl: A benchmarking framework for the Julia...](https://github.com/JuliaCI/BenchmarkTools.jl)**
>
> A benchmarking framework for the Julia language. Contribute to JuliaCI/BenchmarkTools.jl development by creating an account on GitHub.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [September 3, 2018, 8:16am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/5 "2018-09-03T08:16:36Z")

</div>

But BenchmarkTools would not solve our problem, would it? Isn’t the point of that tool to run code multiple times to get stable results? That’s precisely what we’re hoping to avoid.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 3, 2018, 8:22am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/6 "2018-09-03T08:22:33Z")

</div>

Please read its docs, you can run it just one time (ie two times since it will not measure the first one), but IMO that’s pretty pointless for benchmarking.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 3, 2018, 8:24am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/7 "2018-09-03T08:24:41Z")

</div>

To actually be able to catch regressions you’re going to have to run the suite multiple times, in order to average out random noise from differing configurations, machine states, neutrinos flipping a bit, etc… BenchmarkTools is not just the `@btime` and `@benchmark` macros, but provides a bunch of different configurations to support your own test suites. Check the [documentation](https://github.com/JuliaCI/BenchmarkTools.jl/blob/master/doc/manual.md) for a basic introduction into running custom tests. In particular, the part about a [`Trial`](https://github.com/JuliaCI/BenchmarkTools.jl/blob/master/doc/manual.md#trial-and-trialestimate) and what it contains (e.g. number of allocations and allocated memory) is going to be of interest to you. As far as I can tell, if all you’re looking for is the memory used and not the time, running just once should be fine, as the allocations shouldn’t really change.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [September 3, 2018, 9:40am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/8 "2018-09-03T09:40:29Z")

</div>

Hmm, thanks for the responses. I’ve used BenchmarkTools before, and just don’t see how it would help us since in the end its solution to the problem is to run code repeatedly (if even twice), which is what we’re hoping to avoid.

We run the test suite on the build server on every commit, and the build already takes ~10 minutes. What we are looking for are suggestions on how to get _somewhat accurate_ performance measurements for each build – we don’t require nanosecond or single byte precision, but if the memory usage for a test grows by say 2x in a commit, that’s something we’ll want to find out as soon as possible. (And yes, memory usage is fairly stable when testing identical code once, but it’s not reliable at all; tiny irrelevant code changes can change JIT compiling behavior which allocates a lot more or less memory.)

It might just be the nature of Julia with its JIT compiler that what we’re asking for is not possible.

If so, how do other package developers do this?

Faster unit tests, so that they can be benchmarked properly for each build?  
Not running benchmarks as part of the build, so that the benchmarks can be run for longer?  
They don’t benchmark their unit tests?

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 3, 2018, 9:48am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/9 "2018-09-03T09:48:02Z")

</div>

Well, I’m not sure how other developers handle this, but I for one don’t run the full testsuite on every commit. with BenchmarkTools you are able to create test suites for specific parts and just run those - you’ll need some setup though to only run the tests for the code you changed (or what’s affected by the change). Those suites should only run once too, since you don’t care about the time in those parts.

In general though, I don’t think checking for regression of everything on every commit is a good choice, especially because that really limits iteration speed on your code. Your options are probably limited to “run the tests less often” and “write granular testsuites to be run independently”. A combination of both is probably a good choice.

Another idea would be to only benchmark unit tests when you already know they’re slow/use more memory, instead of all the time.

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [September 3, 2018, 10:24am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/10 "2018-09-03T10:24:58Z")

</div>

You might find the readme and the solution of this project useful  
[https://github.com/JuliaCI/BaseBenchmarks.jl](https://github.com/JuliaCI/BaseBenchmarks.jl)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 3, 2018, 10:59am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/11 "2018-09-03T10:59:35Z")

</div>

> [@bennedich](#):
>
> just don’t see how it would help us since in the end its solution to the problem is to run code repeatedly (if even twice), which is what we’re hoping to avoid.

If the problem has a scale (eg number of observations, grid size) that leaves the algorithm invariant, you could run on a small scale to compile, then benchmark on the large one.

---

<div class="post-metadata">

**Author:** ![jonathanBieler](https://avatars.discourse-cdn.com/v4/letter/j/82dd89/32.png) [@jonathanBieler](https://discourse.julialang.org/u/jonathanBieler)\
**Post date:** [September 3, 2018, 1:04pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/12 "2018-09-03T13:04:24Z")

</div>

Basically you want to compile your function before calling it for the first time. Calling `@code_native` seems to compile your function, and you can call `code_native` with a dummy buffer to avoid printing the output. Maybe that doesn’t always work, or there’s a better way to trigger compilation.

```julia
julia> function test()
           A = [x for x=0.0:2]
       end
test (generic function with 1 method)

julia> code_native(IOBuffer(),test,())

julia> @time test()
  0.000030 seconds (6 allocations: 352 bytes)
3-element Array{Float64,1}:
 0.0
 1.0
 2.0

```

---

<div class="post-metadata">

**Author:** ![greg\_plowman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/greg_plowman/32/8100_2.png) [@greg\_plowman](https://discourse.julialang.org/u/greg_plowman)\
**Post date:** [September 3, 2018, 1:11pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/13 "2018-09-03T13:11:51Z")

</div>

> [@jonathanBieler](#):
>
> Basically you want to compile your function before calling it for the first time.

You could use `precompile` for this.

[https://docs.julialang.org/en/v1/base/base/#Base.precompile](https://docs.julialang.org/en/v1/base/base/#Base.precompile)

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 3, 2018, 1:42pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/14 "2018-09-03T13:42:00Z")

</div>

Does it precompile recursively (ie also the functions it calls)?

---

<div class="post-metadata">

**Author:** ![antoine-levitt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/antoine-levitt/32/4008_2.png) [@antoine-levitt](https://discourse.julialang.org/u/antoine-levitt)\
**Post date:** [September 3, 2018, 2:38pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/15 "2018-09-03T14:38:25Z")

</div>

I think that’s impossible in full generality without actually running the code. It may be possible to precompile recursively the calls that are resolved with static dispatch, but that might be problematic also (because it might precompile too much)

---

<div class="post-metadata">

**Author:** ![Tero\_Frondelius](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tero_frondelius/32/7629_2.png) [@Tero\_Frondelius](https://discourse.julialang.org/u/Tero_Frondelius)\
**Post date:** [September 3, 2018, 3:14pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/16 "2018-09-03T15:14:27Z")

</div>

Can you test how much time is needed to run something twice? Is it 20% longer or 99% longer. Depending the result you will need different strategy: 20% case I would just live with and the second case I would look into precompiling.

---

<div class="post-metadata">

**Author:** ![Sukera](https://avatars.discourse-cdn.com/v4/letter/s/ce7236/32.png) [@Sukera](https://discourse.julialang.org/u/Sukera)\
**Post date:** [September 3, 2018, 7:19pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/17 "2018-09-03T19:19:09Z")

</div>

I found some more stuff that might be of interest to you:

[https://docs.julialang.org/en/v1/manual/profile/#Memory-allocation-analysis-1](https://docs.julialang.org/en/v1/manual/profile/#Memory-allocation-analysis-1)

> **[GitHub - JuliaCI/Coverage.jl: Take Julia code coverage and memory allocation...](https://github.com/JuliaCI/Coverage.jl)**
>
> Take Julia code coverage and memory allocation results, do useful things with them - GitHub - JuliaCI/Coverage.jl: Take Julia code coverage and memory allocation results, do useful things with them

Especially Coverage.jl, since single function memory allocation can be tested. This could be integrated into a workflow where changed code gets tested automatically for memory allocation regressions.

---

<div class="post-metadata">

**Author:** ![bennedich](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bennedich/32/4894_2.png) [@bennedich](https://discourse.julialang.org/u/bennedich)\
**Post date:** [September 4, 2018, 8:53am UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/18 "2018-09-04T08:53:15Z")

</div>

Thanks for all the advice! Several good suggestions here, which we’ll look further into.

I am a bit skeptical to relying on code coverage for any part of this, since I find that it’s not working very well in Julia 0.7. I just started [a new topic](https://discourse.julialang.org/t/code-coverage-works-poorly-in-julia-0-7/14522) about this.

---

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [December 22, 2018, 6:49pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/19 "2018-12-22T18:49:53Z")

</div>

> [@bennedich](#):
>
> What we are looking for are suggestions on how to get _somewhat accurate_ performance measurements for each build – we don’t require nanosecond or single byte precision, but if the memory usage for a test grows by say 2x in a commit, that’s something we’ll want to find out as soon as possible

`julia -O, --optimize={0,1,2,3}` Set the optimization level (default 2 if unspecified or 3 if specified as -O)

I don’t know why Julia got slowing, but assume, since there’s a tradeoff for compilation speed and optimization, that the default level is now more aggressive. Maybe lower, e.g. -01 helps?

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [December 22, 2018, 7:13pm UTC](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444/20 "2018-12-22T19:13:23Z")

</div>

> [@Palli](#):
>
> I don’t know why Julia got slowing, but assume, since there’s a tradeoff for compilation speed and optimization, that the default level is now more aggressive. Maybe lower, e.g. -01 helps?

Changing the optimization level seems like a bad idea when you want to test the performance.

[Next page](https://discourse.julialang.org/t/tracking-memory-usage-in-unit-tests-a-lot-worse-in-julia-0-7-than-0-6/14444.md?page=2)
