# Optimizing concatenation of data from large number of files

**URL:** <https://discourse.julialang.org/t/optimizing-concatenation-of-data-from-large-number-of-files/46617>\
**Category:** Performance\
**Created:** [September 14, 2020, 11:58pm UTC](https://discourse.julialang.org/t/optimizing-concatenation-of-data-from-large-number-of-files/46617 "2020-09-14T23:58:27Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![lwhitefox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lwhitefox/32/4363_2.png) [@lwhitefox](https://discourse.julialang.org/u/lwhitefox)\
**Post date:** [September 14, 2020, 11:58pm UTC](https://discourse.julialang.org/t/optimizing-concatenation-of-data-from-large-number-of-files/46617/1 "2020-09-14T23:58:27Z")

</div>

I am trying to create a large data array from data in many input files by

1. reading in the individual data file arrays
2. concatenating a subset of columns to generate a large, single array

The problem is that it is blazingly fast for the first, say, 100 files (depends on how I try to concatenate) but then each subsequent iteration slows down dramatically. So, I’m obviously doing something very inefficient - I’d assume trying to use `vcat()` would be slow because it would keep allocating to pad the current array. For this reason, I instead allocate the entire final array up front and try to fill in the data as it is read from individual files.

I include an MWE example below - because all test arrays are small, it is very quick - but I’m hoping someone will see what is bad about the approach as I need to to run quickly for much larger arrays (individual arrays read in are on the order to 15000 x 1000).

```julia
using Glob

n_columns = 100
n_rows = 100
# create fake data
for i in 1:100
    d = rand(n_rows,n_columns)
    write("file_$(i).bin",d)
end

idx_cols_keep = 10:20

data_in = Array{Float32,2}(undef,n_rows,n_columns);
# yes order will be "wrong" comapred to indices but this is just an MWE to demo the idea
filelist = glob("*.bin",".");

nrows_total = n_rows*length(filelist)

data_total = Array{Float32,2}(undef,nrows_total,length(idx_cols_keep));

i0 = 1
nmax = length(filelist)
for i in 1:nmax
    read!(filelist[i],data_in)
    data_total[i0:i0+n_rows-1,:] .= data_in[:,idx_cols_keep]
    i0 += n_rows
end

```

---

<div class="post-metadata">

**Author:** ![baggepinnen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/baggepinnen/32/693_2.png) [@baggepinnen](https://discourse.julialang.org/u/baggepinnen)\
**Post date:** [September 15, 2020, 4:59am UTC](https://discourse.julialang.org/t/optimizing-concatenation-of-data-from-large-number-of-files/46617/2 "2020-09-15T04:59:11Z")

</div>

I would suggest trying to put the code in a function as opposed to in global scope. The compiler will think much harder about the code that way, and can possibly optimize memory usage better.  
See [Performance Tips · The Julia Language](https://docs.julialang.org/en/v1/manual/performance-tips/)

---

<div class="post-metadata">

**Author:** ![lwhitefox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lwhitefox/32/4363_2.png) [@lwhitefox](https://discourse.julialang.org/u/lwhitefox)\
**Post date:** [September 15, 2020, 5:05pm UTC](https://discourse.julialang.org/t/optimizing-concatenation-of-data-from-large-number-of-files/46617/3 "2020-09-15T17:05:26Z")

</div>

You were right - I was stupidly assuming that this was a simple enough operation that there was no need to create a function even though I “knew” this rule.

Sigh.
