# Help with Malaria dataset

**URL:** <https://discourse.julialang.org/t/help-with-malaria-dataset/25004>\
**Category:** New to Julia\
**Created:** [June 6, 2019, 10:36am UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004 "2019-06-06T10:36:35Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Deanodirector](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deanodirector/32/37411_2.png) [@Deanodirector](https://discourse.julialang.org/u/Deanodirector)\
**Post date:** [June 6, 2019, 10:36am UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/1 "2019-06-06T10:36:35Z")

</div>

Hello, just starting out with Julia here.

I’ve downloaded the Malaria dataset from kaggle and I have two folders of images, one parasitized cells and the other normal cells. How do I go about loading all these images into an array?

I could do a for loop but I don’t know the command to get the number of image files in a folder. The file names aren’t consistent either.

Many thanks

---

<div class="post-metadata">

**Author:** ![Ayushk4](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ayushk4/32/8354_2.png) [@Ayushk4](https://discourse.julialang.org/u/Ayushk4)\
**Post date:** [June 6, 2019, 10:58am UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/2 "2019-06-06T10:58:52Z")

</div>

You could obtain all the files and directories: `readdir(dir_path)`.  
Then check if an element is a file or a directory: `isfile(file_path)`.  
Then for each element, you could check for the file extension: `endswith(file_name, suffix)`.

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [June 6, 2019, 11:17am UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/3 "2019-06-06T11:17:39Z")

</div>

Deano have you looked at Metalhead [https://github.com/FluxML/Metalhead.jl](https://github.com/FluxML/Metalhead.jl)

---

<div class="post-metadata">

**Author:** ![Deanodirector](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deanodirector/32/37411_2.png) [@Deanodirector](https://discourse.julialang.org/u/Deanodirector)\
**Post date:** [June 6, 2019, 12:47pm UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/4 "2019-06-06T12:47:12Z")

</div>

thanks! I’m getting error no method matching joinpath() when I try to filter the files

```julia
filter(isfile(), readdir())

```

I’m guessing I have to format the strings somehow?

---

<div class="post-metadata">

**Author:** ![ericphanson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ericphanson/32/215186_2.png) [@ericphanson](https://discourse.julialang.org/u/ericphanson)\
**Post date:** [June 6, 2019, 1:07pm UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/5 "2019-06-06T13:07:36Z")

</div>

Use

```julia
filter(isfile, readdir())

```

instead-- you want to pass the function `isfile` to `filter`, whereas in your version, you’re calling `isfile` (with no arguments), and then passing the result to `filter`.

---

<div class="post-metadata">

**Author:** ![Deanodirector](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deanodirector/32/37411_2.png) [@Deanodirector](https://discourse.julialang.org/u/Deanodirector)\
**Post date:** [June 6, 2019, 3:32pm UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/6 "2019-06-06T15:32:38Z")

</div>

I should have seen that. thank you!

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [June 6, 2019, 4:31pm UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/7 "2019-06-06T16:31:50Z")

</div>

[Glob.jl](https://github.com/vtjnash/Glob.jl) is useful for this sort of thing:

```julia
using Glob
readdir(glob"*.jpg", "/path/to/image/dir")
```

---

<div class="post-metadata">

**Author:** ![Deanodirector](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/deanodirector/32/37411_2.png) [@Deanodirector](https://discourse.julialang.org/u/Deanodirector)\
**Post date:** [June 7, 2019, 10:11am UTC](https://discourse.julialang.org/t/help-with-malaria-dataset/25004/8 "2019-06-07T10:11:18Z")

</div>

Brilliant responses thank you!

I now have

```julia
parapath = "../Kaggle Datasets/cell-images-for-detecting-malaria/cell_images/Parasitized/"
parafiles = readdir(glob"*.png", parapath)
paraimages = hcat(reshape(map(load,parafiles),:))

```

and the same for the uninfected images. I then use

```julia
X = vcat(paraimages,unimages)

```

to get my input array. Is this the standard input type for flux?
