# Val{N} + LinearIndices Causes Massive Compile-Time Unrolling

**URL:** <https://discourse.julialang.org/t/val-n-linearindices-causes-massive-compile-time-unrolling/128248>\
**Category:** GPU\
**Tags:** question, kernelabstractions, reflection, lowering\
**Created:** [April 20, 2025, 1:27pm UTC](https://discourse.julialang.org/t/val-n-linearindices-causes-massive-compile-time-unrolling/128248 "2025-04-20T13:27:55Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![0samuraiE](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/0samuraie/32/209825_2.png) [@0samuraiE](https://discourse.julialang.org/u/0samuraiE)\
**Post date:** [April 20, 2025, 1:27pm UTC](https://discourse.julialang.org/t/val-n-linearindices-causes-massive-compile-time-unrolling/128248/1 "2025-04-20T13:27:55Z")

</div>

I’m seeing strange behavior when using `@code_warntype` on a function that takes a `Val{N}` parameter and creates a `LinearIndices` object with size `(N, N, N)`. Here’s a minimal example:

```julia
function test(::Val{N}) where {N}
    LinearIndices((1:N, 1:N, 1:N))
end
@code_warntype test(Val(100))

```

The output includes a massive list of literal integers:

```julia
996900 997000 997100 997200 ... 1000000])
└── return %9

```

This was surprising and prompted me to investigate further, because I encountered a serious performance issue when using similar code in a GPU kernel via KernelAbstractions.jl. Specifically, compiling the kernel resulted in:

```julia
588.306677 seconds (13.85 M CPU allocations: 2.513 GiB, 0.05% gc time), 0.00% GPU memmgmt time

```

It seems that passing a large N as a type parameter (Val{N}) causes excessive unrolling or compile-time expansion of index structures. My questions are:

Is this behavior expected when using Val{N} with large N in combination with LinearIndices?

Is there a recommended way to avoid this kind of compile-time explosion while maintaining type stability?

Are there best practices for using Val-based parameters safely in GPU kernels with KernelAbstractions?

Any insight would be appreciated.

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [April 20, 2025, 2:12pm UTC](https://discourse.julialang.org/t/val-n-linearindices-causes-massive-compile-time-unrolling/128248/2 "2025-04-20T14:12:48Z")

</div>

> [@0samuraiE](#):
>
> The output includes a massive list of literal integers:

On v1.11 and below, this displayed form is used because the 2-argument `show` for `LinearIndices` is not specialized, and it falls back to the default `show` for `AbstractArray`s. A specialized method is now added on the upcoming v1.12, so we display something a bit more meaningful:

```julia
julia> @code_warntype test(Val(6))
MethodInstance for test(::Val{6})
  from test(::Val{N}) where N @ Main REPL[1]:1
Static Parameters
  N = 6
Arguments
  #self#::Core.Const(Main.test)
  _::Core.Const(Val{6}())
Body::LinearIndices{3, Tuple{UnitRange{Int64}, UnitRange{Int64}, UnitRange{Int64}}}
1 ─ %1 = Main.LinearIndices::Core.Const(LinearIndices)
│ %2 = Main.:(:)::Core.Const(Colon())
│ %3 = $(Expr(:static_parameter, 1))::Core.Const(6)
│ %4 = (%2)(1, %3)::Core.Const(1:6)
│ %5 = Main.:(:)::Core.Const(Colon())
│ %6 = $(Expr(:static_parameter, 1))::Core.Const(6)
│ %7 = (%5)(1, %6)::Core.Const(1:6)
│ %8 = Main.:(:)::Core.Const(Colon())
│ %9 = $(Expr(:static_parameter, 1))::Core.Const(6)
│ %10 = (%8)(1, %9)::Core.Const(1:6)
│ %11 = Core.tuple(%4, %7, %10)::Core.Const((1:6, 1:6, 1:6))
│ %12 = (%1)(%11)::Core.Const(LinearIndices((1:6, 1:6, 1:6)))
└── return %12

```

I doubt this is where your performance issue arises from, since the struct simply contains three `UnitRange`s.

---

<div class="post-metadata">

**Author:** ![0samuraiE](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/0samuraie/32/209825_2.png) [@0samuraiE](https://discourse.julialang.org/u/0samuraiE)\
**Post date:** [April 20, 2025, 2:47pm UTC](https://discourse.julialang.org/t/val-n-linearindices-causes-massive-compile-time-unrolling/128248/3 "2025-04-20T14:47:46Z")

</div>

Thanks again. Here’s a concise summary of what I observed:

- Removing `LinearIndices` from the GPU kernel significantly improved compile-time performance in my app.

- The slowdown only occurs on the **first execution** , which strongly suggests a compile-time issue rather than runtime inefficiency.

- When axes=(1:64,1:64,1:64), first execution is not a problem, but when axes=(1:256, 1:256, 1:256) first execution is slow.

- Based on this, I suspect that `LinearIndices` may be triggering excessive specialization or compile-time computation when used in CUDA kernels, perhaps due to the complex type structure it introduces. (Though the idea that this causes large memory pressure during compilation is just my own speculation.)

In my case, I didn’t strictly need `LinearIndices`, so switching to a simple linear range solved the issue. However, I’d still appreciate any insight into:

- Why `LinearIndices` causes such heavy compile-time behavior in GPU code,
- And whether there’s a better way to handle multi-dimensional indexing efficiently in kernels.

And maybe this is more related to GPU compilation specifically, so I’ll move this topic to a more appropriate category.  
Any advice or internal explanation would be very helpful.
