# Data spread over multiple files

**URL:** <https://discourse.julialang.org/t/data-spread-over-multiple-files/15464>\
**Category:** Machine Learning\
**Created:** [September 25, 2018, 12:10pm UTC](https://discourse.julialang.org/t/data-spread-over-multiple-files/15464 "2018-09-25T12:10:16Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [September 25, 2018, 12:10pm UTC](https://discourse.julialang.org/t/data-spread-over-multiple-files/15464/1 "2018-09-25T12:10:16Z")

</div>

Hi,

I would like to ask, if someone has dealt in principle manner with problem, where the data for machine learning is spread over several files.

I have once defined an iterator, which takes loadfunction and list of files as an input and outputs minibatches with a fixed number of observations irrespectively the number of samples in each file.

The code was written for julia-0.6 and I have not updated it yet to 0.7 / 1.0. I would like to know, if someone did something similar or if there is a desire for such a thing.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 25, 2018, 1:26pm UTC](https://discourse.julialang.org/t/data-spread-over-multiple-files/15464/2 "2018-09-25T13:26:18Z")

</div>

Hard to say more without an MWE, but if you have iterators for individual files, you can use [`Base.Iterators.flatten`](https://docs.julialang.org/en/stable/base/iterators/#Base.Iterators.flatten) to concatenate them, and then optionally [`Base.Iterators.partition`](https://docs.julialang.org/en/stable/base/iterators/#Base.Iterators.partition) or similar for small batches.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [September 25, 2018, 6:37pm UTC](https://discourse.julialang.org/t/data-spread-over-multiple-files/15464/3 "2018-09-25T18:37:59Z")

</div>

Hi Tamas,

thank you very much for the answer.  
Following is a test I have written:

```julia
using DataIterators: FileIterator

d = Dict("a" => [1 2 3 4 5], 
	"b" => [6 7 8 9 10], 
	"c" => [11 12 13 14])
loadfun(f) = (println("reading ",f); d[f])

files = ["a", "b", "c"]

begin 
	iter = FileIterator(loadfun, files, 2)
	nxt = iterate(iter)
	i = 0
	while nxt !== nothing && i < 10
		(x, state) = nxt
		println(x)
		nxt = iterate(iter, state)
		i += 1
	end
end

```

The output should be:

```julia
reading a
[1 2]
[3 4]
reading b
[5 6]
[7 8]
[9 10]
reading c
[11 12]
[13 14]

```

I think that flatten does not do the job here, since I need the data to be loaded lazily when needed. I have a first proof of concept implementation, certainly far from being perfect.

---

<div class="post-metadata">

**Author:** ![Tomas\_Pevny](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tomas_pevny/32/25466_2.png) [@Tomas\_Pevny](https://discourse.julialang.org/u/Tomas_Pevny)\
**Post date:** [September 26, 2018, 4:55am UTC](https://discourse.julialang.org/t/data-spread-over-multiple-files/15464/4 "2018-09-26T04:55:45Z")

</div>

If someone is interested in this problem, here is a first shot on this

> **[GitHub - pevnak/DataIterators.jl](https://github.com/pevnak/DataIterators.jl)**
>
> Contribute to pevnak/DataIterators.jl development by creating an account on GitHub.
