# Using LinuxPerf for small functions

**URL:** <https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410>\
**Category:** Performance\
**Tags:** question, package\
**Created:** [May 2, 2021, 5:03am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410 "2021-05-02T05:03:04Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![shmiggles](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shmiggles/32/19871_2.png) [@shmiggles](https://discourse.julialang.org/u/shmiggles)\
**Post date:** [May 2, 2021, 5:03am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/1 "2021-05-02T05:03:04Z")

</div>

LinuxPerf.jl wraps the `perf_event_open` Linux syscall. But using it for small functions gives ridiculous results. The following example reports over 12000 clock cycles and 2000 memory fetches to compute 1+1:

```julia
using LinuxPerf
@measure 1+1

```

Dumb question: what is running other than the execution of 1+1 for `perf_event_open` to report so many events? Expression parsing?

---

<div class="post-metadata">

**Author:** ![carstenbauer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/carstenbauer/32/4981_2.png) [@carstenbauer](https://discourse.julialang.org/u/carstenbauer)\
**Post date:** [May 2, 2021, 6:52am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/2 "2021-05-02T06:52:00Z")

</div>

@vchuravy ☝

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [May 2, 2021, 12:02pm UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/3 "2021-05-02T12:02:47Z")

</div>

First, run `@measure` twice, and in a function: `f() = @measure 1+1; f(); f()`. It does not run the expression multiple times and average over runs like `BenchmarkTools.@btime` does, so the first time you run this, you’re getting compilation time. And running it at the global scope includes some penalty from running code at the toplevel.

Second, this function is far too small to be effectively measured by the perf subsystem. You’re basically doing a syscall, performing 1-2 instructions, and then immediately performing another syscall. Performing each of those syscalls requires a non-negligible number of instructions (which perf will count), both in Julia and in the kernel. You could try running this in a repeated loop for some number of iterations, but you’d still be picking up loop overhead (at least 2 extra instructions) after calculating the average result.

---

<div class="post-metadata">

**Author:** ![shmiggles](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shmiggles/32/19871_2.png) [@shmiggles](https://discourse.julialang.org/u/shmiggles)\
**Post date:** [May 6, 2021, 4:32am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/4 "2021-05-06T04:32:49Z")

</div>

Running @measurement multiple times still gives me thousands of cycles (6000+, though quite variable across runs) and 2270 instructions.

Doing an equivalent benchmark in C gives me ~60 cycles and 13 instructions.

---

<div class="post-metadata">

**Author:** ![Elrod](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/elrod/32/22461_2.png) [@Elrod](https://discourse.julialang.org/u/Elrod)\
**Post date:** [May 6, 2021, 5:01am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/5 "2021-05-06T05:01:58Z")

</div>

I always define

```julia
function foreachf(f::F, N, args::Vararg{Any,A}) where {F,A}
    foreach(_ -> f(args...), 1:N)
end

```

So that it calls `f(args...)` a total of `N` times.  
However, you’ll have to make sure the compiler doesn’t defeat the benchmark, like it does for `+(::Int,::Int)`.

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [May 6, 2021, 4:31pm UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/6 "2021-05-06T16:31:23Z")

</div>

> [@shmiggles](#):
>
> Doing an equivalent benchmark in C gives me ~60 cycles and 13 instructions.

Can you show the code for this C benchmark?

---

<div class="post-metadata">

**Author:** ![shmiggles](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/shmiggles/32/19871_2.png) [@shmiggles](https://discourse.julialang.org/u/shmiggles)\
**Post date:** [May 14, 2021, 4:33am UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/7 "2021-05-14T04:33:19Z")

</div>

Sure.

```nohighlight
#include <stdlib.h>
#include <stdio.h>
#include <unistd.h>
#include <string.h>
#include <sys/ioctl.h>
#include <linux/perf_event.h>
#include <asm/unistd.h>

static long
perf_event_open(struct perf_event_attr *hw_event, pid_t pid,
                int cpu, int group_fd, unsigned long flags)
{
    int ret;

    ret = syscall(__NR_perf_event_open, hw_event, pid, cpu,
                   group_fd, flags);
    return ret;
}

int
main(int argc, char **argv)
{
    struct perf_event_attr pe;
    long long count;
    int fd;

    memset(&pe, 0, sizeof(struct perf_event_attr));
    pe.type = PERF_TYPE_HARDWARE;
    pe.size = sizeof(struct perf_event_attr);
    pe.config = PERF_COUNT_HW_CPU_CYCLES;
    pe.disabled = 1;
    pe.exclude_kernel = 1;
    pe.exclude_hv = 1;

    fd = perf_event_open(&pe, 0, -1, -1, 0);
    ioctl(fd, PERF_EVENT_IOC_RESET, 0);
    ioctl(fd, PERF_EVENT_IOC_ENABLE, 0);
    int x = 1+1;
    ioctl(fd, PERF_EVENT_IOC_DISABLE, 0);
    read(fd, &count, sizeof(long long));

    printf("Used %lld cyles\n", count);

    close(fd);
    return x-2;
}

```

```julia
gcc perftest.c -o perftest; chmod +x perftest; ./perftest

```

---

<div class="post-metadata">

**Author:** ![jpsamaroo](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jpsamaroo/32/46804_2.png) [@jpsamaroo](https://discourse.julialang.org/u/jpsamaroo)\
**Post date:** [May 15, 2021, 2:39pm UTC](https://discourse.julialang.org/t/using-linuxperf-for-small-functions/60410/8 "2021-05-15T14:39:09Z")

</div>

```julia
using LinuxPerf

bench = make_bench([LinuxPerf.EventType(:hw, :cycles)])
function f(bench, x)
  enable!(bench)
  x = x+1
  disable!(bench)
  x
end
f(bench, x)
reset!(bench)
f(bench, x)
@show counters(bench)

```

I get 96 cycles from the above. So probably you just need to change the default bench (which defaults to `reasonable_defaults`, which are actually a lot of metrics the kernel needs to collect and process) and make sure to only measure within a function.
