# Normal vs broadcasted slice assignment

**URL:** <https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288>\
**Category:** General Usage\
**Created:** [February 16, 2024, 10:24am UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288 "2024-02-16T10:24:02Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sijo](https://avatars.discourse-cdn.com/v4/letter/s/da6949/32.png) [@sijo](https://discourse.julialang.org/u/sijo)\
**Post date:** [February 16, 2024, 10:24am UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/1 "2024-02-16T10:24:02Z")

</div>

This must have been discussed a dozen times but I couldn’t find a thread about this precise issue:

```julia
using BenchmarkTools

f1(v, x) = v[1:length(x)] = x
f2(v, x) = v[1:length(x)] .= x

julia> @btime f1($(rand(1000)), $(rand(100)));
  10.268 ns (0 allocations: 0 bytes)

julia> @btime f2($(rand(1000)), $(rand(100)));
  25.415 ns (0 allocations: 0 bytes)

```

Is this expected? Does it have to do with unaliasing? I wonder what’s causing the slowdown exactly since there’s 0 allocation in both cases.

---

<div class="post-metadata">

**Author:** ![mbauman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mbauman/32/31082_2.png) [@mbauman](https://discourse.julialang.org/u/mbauman)\
**Post date:** [February 16, 2024, 5:49pm UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/2 "2024-02-16T17:49:28Z")

</div>

No, I think this is just the difference between a highly specialized `memcpy` that hits the `Vector`’s memory directly and a hand-written `for` loop that works with all abstract arrays.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [February 16, 2024, 6:19pm UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/3 "2024-02-16T18:19:01Z")

</div>

These both should turn into memcpy. IMO this is unexpected.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [February 16, 2024, 6:34pm UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/4 "2024-02-16T18:34:55Z")

</div>

So the difference is that `f1` turns into a `copyto!(view(a, 1:100), b)` while `f2` turns into a `setindex!`. So the problem is just that we don’t have an optimized method for copying a `view` of an `Array` to another `Array`.

---

<div class="post-metadata">

**Author:** ![danielwe](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/danielwe/32/35657_2.png) [@danielwe](https://discourse.julialang.org/u/danielwe)\
**Post date:** [February 16, 2024, 7:25pm UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/5 "2024-02-16T19:25:43Z")

</div>

Another difference is that `f2` returns a view of `v`, while `f1` returns `x`. Changing both functions to return `nothing` improves the performance of `f2` somewhat, although not enough to make up the difference.

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [February 16, 2024, 8:10pm UTC](https://discourse.julialang.org/t/normal-vs-broadcasted-slice-assignment/110288/6 "2024-02-16T20:10:26Z")

</div>

In

> [@Why is copying using a loop is much slower than \`copy\` for large arrays?](https://discourse.julialang.org/t/why-is-copying-using-a-loop-is-much-slower-than-copy-for-large-arrays/90633/2):
>
> A possible explaination : The builtin-in copy simply calls a builtin function jl\_array\_copy, which in turn calls a builtin array constructor and memcpy function. Your simple implementation cannot beat C’s highly optimized memcpy (memcpy is really smart and it can sometimes even utilize special feature offered by operating system). The case of small array may be related to alignment problem or overhead in memcpy, since you copy directly and LLVM knows more information, some logic in memcpy can b…

Chris Elrod had suggested that the differences arise from non-temporal stores, and that LoopVectorization provides a Julia equivalent.
