# Avoid LLVM setjmp bug

**URL:** <https://discourse.julialang.org/t/avoid-llvm-setjmp-bug/1140>\
**Category:** Internals & Design\
**Created:** [December 24, 2016, 7:43pm UTC](https://discourse.julialang.org/t/avoid-llvm-setjmp-bug/1140 "2016-12-24T19:43:11Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![jameson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jameson/32/23_2.png) [@jameson](https://discourse.julialang.org/u/jameson)\
**Post date:** [December 24, 2016, 7:43pm UTC](https://discourse.julialang.org/t/avoid-llvm-setjmp-bug/1140/1 "2016-12-24T19:43:11Z")

</div>

Currently, we can’t reliably run LLVM optimizations on any function with a try/catch since it mis-optimizes it. This is captured in the following issue:

> <https://github.com/JuliaLang/julia/issues/17288>
>
> One of my packages (CUDAdrv) has recently started failing on julia master, with …a segfault in \`typemap.c\`. I've bisected this issue to e2bd1298732f36465fbd5d112d959fe1de052c7c (all backtraces and line numbers below are on that commit's tree). I'm not sure where to start debugging this, so I'm at least reporting it here already.
> 
> This causes a segfault in \`jl\_typemap\_level\_assoc\_exact\`:
> 
> \`\`\`
> signal (11): Segmentation fault
> while loading CUDAdrv/test/core.jl, in expression starting on line 172
> sig\_match\_fast at src/gf.c:1707
> jl\_apply\_generic at src/gf.c:1886
> Type at CUDAdrv/src/module.jl:67
> unknown function (ip: 0x7f6198ad561e)
> jl\_call\_method\_internal at src/julia\_internal.h:92
> jl\_apply\_generic at src/gf.c:1931
> do\_call at src/interpreter.c:65
> eval at src/interpreter.c:188
> eval\_body at src/interpreter.c:469
> eval\_body at src/interpreter.c:515
> jl\_interpret\_call at src/interpreter.c:573
> jl\_interpret\_toplevel\_thunk at src/interpreter.c:580
> jl\_toplevel\_eval\_flex at src/toplevel.c:543
> jl\_parse\_eval\_all at src/ast.c:700
> jl\_load at src/toplevel.c:566
> jl\_load\_ at src/toplevel.c:575
> include\_from\_node1 at ./loading.jl:426
> unknown function (ip: 0x7f639ef3716c)
> jl\_call\_method\_internal at src/julia\_internal.h:92
> jl\_apply\_generic at src/gf.c:1931
> do\_call at src/interpreter.c:65
> eval at src/interpreter.c:188
> jl\_interpret\_toplevel\_expr at src/interpreter.c:31
> jl\_toplevel\_eval\_flex at src/toplevel.c:529
> jl\_parse\_eval\_all at src/ast.c:700
> jl\_load at src/toplevel.c:566
> jl\_load\_ at src/toplevel.c:575
> include\_from\_node1 at ./loading.jl:426
> unknown function (ip: 0x7f639ef3716c)
> jl\_call\_method\_internal at src/julia\_internal.h:92
> jl\_apply\_generic at src/gf.c:1931
> process\_options at ./client.jl:266
> \_start at ./client.jl:322
> unknown function (ip: 0x7f639ef75124)
> jl\_call\_method\_internal at src/julia\_internal.h:92
> jl\_apply\_generic at src/gf.c:1931
> jl\_apply at ui/../src/julia.h:1396
> true\_main at ui/repl.c:546
> main at ui/repl.c:674
> unknown function (ip: 0x7f63a5013740)
> unknown function (ip: 0x401818)
> Allocations: 1748036 (Pool: 1746766; Other: 1270); GC: 4
> Allocations: 1748036 (Pool: 1746766; Other: 1270); GC: 4
> \`\`\`
> 
> Running in GDB makes it segfault somewhere else, but I assume due to the same problem (\`jl\_typeof(NULL)\`):
> 
> \`\`\`
> Thread 1 "julia" received signal SIGSEGV, Segmentation fault.
> 0x00007ffff76c9407 in jl\_typemap\_level\_assoc\_exact (cache=0x7ffdf1802950, args=0x7fffffffae90, n=3, offs=1 '\\001') at src/typemap.c:788
> 788 jl\_value\_t \*ty = (jl\_value\_t\*)jl\_typeof(a1);
> 
> 
> (gdb) l
> 783 
> 784 jl\_typemap\_entry\_t \*jl\_typemap\_level\_assoc\_exact(jl\_typemap\_level\_t \*cache, jl\_value\_t \*\*args, size\_t n, int8\_t offs)
> 785 {
> 786 if (n \> offs) {
> 787 jl\_value\_t \*a1 = args\[offs\];
> 788 jl\_value\_t \*ty = (jl\_value\_t\*)jl\_typeof(a1);
> 789 assert(jl\_is\_datatype(ty));
> 790 if (ty == (jl\_value\_t\*)jl\_datatype\_type && cache-\>targ != (void\*)jl\_nothing) {
> 791 union jl\_typemap\_t ml\_or\_cache = mtcache\_hash\_lookup(cache-\>targ, a1, 1, offs);
> 792 jl\_typemap\_entry\_t \*ml = jl\_typemap\_assoc\_exact(ml\_or\_cache, args, n, offs+1);
> 
> (gdb) call jl\_(args\[0\])
> Base.#==()
> (gdb) p args\[1\]
> $3 = (jl\_value\_t \*) 0x0
> (gdb) call jl\_(args\[2\])
> CUDAdrv.CuError(code=209, info=Base.Nullable{String}(isnull=true, value=#\<null\>))
> \`\`\`
> 
> ... with this comparison (against 209 == \`CUDAdrv.ERROR\_NO\_BINARY\_FOR\_GPU\`) originating from:
> 
> \`\`\` julia
> try
> @apicall(:cuModuleLoadDataEx,
> (Ptr{CuModule\_t}, Ptr{Cchar}, Cuint, Ref{CUjit\_option}, Ref{Ptr{Void}}),
> handle\_ref, data, length(optionKeys), optionKeys, optionValues)
> catch err
> (err == ERROR\_NO\_BINARY\_FOR\_GPU || err == ERROR\_INVALID\_IMAGE) || rethrow(err)
> options = decode(optionKeys, optionValues)
> rethrow(CuError(err.code, options\[ERROR\_LOG\_BUFFER\]))
> end
> \`\`\`
> 
> I've not been able to reduce the test case, as reduced versions did not reliably trigger the segfault on all my systems anymore, while the full CUDAdrv test suite does. I've tested on two Linux64 systems (one Debian 8, one Arch), with fresh builds without any Makefile flags.
> 
> @yuyichao any ideas what might be causing this, or where to look for clues?

Codegen bugs are always really nasty, since they can be so unpredictable. Today I realized we may be able to change our codegen representation to avoid this case with minimal effort! If codegen always out-lined the body of try/catch code, it should no longer be able to generate bad code. I think this may even let it perform better optimizations than currently. The key here is that the return path from `setjmp` needs to avoid referencing any state.

In pseudo-C syntax, this would mean codegen would take the Julia function:

```julia
function f()
  setup
  try
    try-body
  catch
    catch-body
  end
  rest-of-function
end

```

And emit something of the form:

```cpp
jl_value_t *f(args, …) { /* the actual function */
  alloca /* local state */
  <setup> /* user code */
  switch (f_trycatch(&alloca)) {
  case 0: /* normal fall-through */
    break;
  case 1: /* user code */
    <catch-body>
    break;
  default:
    abort(); /* corrupted codegen */
  }
  <rest-of-function> /* user code */
}
int f_trycatch(struct* alloca) {
  /* return value describes control flow continuation path */
  /* all other local state (including a gc-frame)
  /* is packaged into the alloca struct that is passed as a pointer argument */
  if (int control-flow = setjmp(&alloca->jmpbuf)) /* Expr(:enter) */
    return control-flow;
  <try-body> /* user code */
  return 0; /* Expr(:leave) */
}

```

Since there’s already a couple function calls on this code path, I don’t think the addition of the extra local `jmp` statement will impact performance. While the removal of `volatile` load / store may actually permit improved codegen optimizations (so, net benefit).

I think this should work, but having a second person review this concept is always good, so I’m posting here for review and other suggestions. 🙂
