# Improvement on appending data using for loop

**URL:** https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686
**Category:** Performance
**Created:** [March 7, 2021, 4:42pm UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686 "2021-03-07T16:42:51Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![FlatWhite17](https://avatars.discourse-cdn.com/v4/letter/f/e79b87/32.png) [@FlatWhite17](https://discourse.julialang.org/u/FlatWhite17)
#### Post date: [March 7, 2021, 4:42pm UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686/1 "2021-03-07T16:42:51Z")

</div>

Hi there,

I’m from a non-CS background and new to Julia. My current project involves a matching process where for each treated unit (D=1), I need to find units in the donor space (D=0) that have variance of their outcome distance across a pre-period (before the treated unit gets treated) smaller than a threshold (e.g. var(y\_j - y\_i \< 0.005).  
My codes are posted below, it works out but very slow. Basically I did the above matching for one treated unit, and it returns a DataFrame called “matched” with the treated unit and all matched donor units. Then I write the below function to do the same matching process on other treated units and keep appending result to the “matched” DataFrame. In total I have 3644 treated units. I checked my CPU usage was only around 20%, though I’m not sure if this process can use multithreading.  
Any advice on how to improve upon it is appreciated, thanks!

 ![Capture](https://global.discourse-cdn.com/julialang/original/3X/9/8/9821656352b828fab42245e4545f0dc7f7fb8e75.png)

---

<div class="post-metadata">

### Author: ![jamblejoe](https://avatars.discourse-cdn.com/v4/letter/j/ee7513/32.png) [@jamblejoe](https://discourse.julialang.org/u/jamblejoe)
#### Post date: [March 14, 2021, 10:46am UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686/2 "2021-03-14T10:46:41Z")

</div>

Hi @FlatWhite17 , welcome to the Julia discourse.  
Some advice for future posts:

- Dont post pictures of your code, but copy it into the code environment ``` ```. When asking to make something more performant people might want to quickly check your code snippets on their computer and do not want to hand copy code.
- Make it as easy as possible for other people to read, understand and execute your code snippets on their machines. Therefore, reduce your snippet as much as possible and avoid variables you did not define in your post. The best would be to provide a minimal working example, people can quickly copy paste into a Julia session.

You already have divided your function `iter_append` into parts marked by comments. Break down your function into smaller parts and see which one is consuming the most time/memory. You can do that quick and dirty with `@time begin ... end`. If you have found the lines which consume the most time you can benchmark them separately, best with the package [`BenchmarkTools`](https://github.com/JuliaCI/BenchmarkTools.jl) and `@btime` or `@benchmark`. This helps you write minimal working examples as well. Alternatively, you can use a profiler, e.g. [`Profile`](https://docs.julialang.org/en/v1/manual/profile/).

---

<div class="post-metadata">

### Author: ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)
#### Post date: [March 14, 2021, 2:05pm UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686/3 "2021-03-14T14:05:36Z")

</div>

In the first line:

```julia
tr_all = filter!(row -> !ismissing(fac_treated[!, outcome]), fac_treated)

```

You take a `row` argument that you use for nothing. Are you sure this is what you want?

---

<div class="post-metadata">

### Author: ![dmolina](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dmolina/32/5246_2.png) [@dmolina](https://discourse.julialang.org/u/dmolina)
#### Post date: [March 14, 2021, 4:07pm UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686/4 "2021-03-14T16:07:10Z")

</div>

Following @Henrique_Becker, I suggest to use the optional column to dropmissing:

```julia
tr_all = dropmissing(fac_treated, :outcome)

```

you’re code do the operation for each row, so it is very expensive.  
I suggest you to review the DataFrame documentation. For instance, you can obtain the mean (be careful, I did not check it):

```julia
base_yr = combine(filter(:ACCPT_ID => ==(fac), tr_all), :T0 => mean)[1,1]

```

---

<div class="post-metadata">

### Author: ![FlatWhite17](https://avatars.discourse-cdn.com/v4/letter/f/e79b87/32.png) [@FlatWhite17](https://discourse.julialang.org/u/FlatWhite17)
#### Post date: [March 14, 2021, 4:26pm UTC](https://discourse.julialang.org/t/improvement-on-appending-data-using-for-loop/56686/5 "2021-03-14T16:26:22Z")

</div>

Thanks! @dmolina! That’s very helpful.

I followed your suggestion and also saved dataframe from each iteration in an array and then use vcat( ) to combine them (instead of append in each iteration). The speed is up a lot.
