# Get JuliaDB.loadtable to parse all columns in CSVs as String

**URL:** <https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [June 2, 2020, 3:40pm UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620 "2020-06-02T15:40:52Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon64288406](https://avatars.discourse-cdn.com/v4/letter/a/da6949/32.png) [@anon64288406](https://discourse.julialang.org/u/anon64288406)\
**Post date:** [June 2, 2020, 3:40pm UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/1 "2020-06-02T15:40:52Z")

</div>

I’m using Julia 1.4.

I want to use JuliaDB.jl, specifically, to read a bunch of CSVs and combine them into one big DataFrame. Here’s the issue: When reading the CSVs, I want all columns to be parsed as String. The number of columns in each CSV differs.

Here’s what I’ve tried:

```julia
using CSV
using DataFrames # just for creating the example DataFrames
using JuliaDB

df1 = DataFrame(
    [['a', 'b', 'c'], [1, 2, 3]],
    ["name", "id"]
)
df2 = DataFrame(
    [['d', 'e', 'f'], [4, 5, 6], [11, 22, 33]],
    ["name", "id", "other"]
)

# For simplicity, I will read just two CSVs, but imagine 20+.
#
# Assume these CSVs are the only files returned by `readdir()`
# below.
CSV.write("df1.csv", df1)
CSV.write("df2.csv", df2)

# This works only if each CSV has the same number of columns with
# the exact same name. But I need it to work for CSVs with
# differing numbers of columns and column names. Also, this gets
# unwieldy if there are many columns.
df = loadtable(readdir(); colparsers=Dict(:name=>String, :id=>String))

# This doesn't work
df = loadtable(readdir(); colparsers=String)
# MethodError: no method matching iterate(::Type{String})

```

Here’s how I’d do it in R:

```r
library(purrr) # Need dplyr installed for `map_dfr()` to work

# Assume list.files() returns just the two above-specified CSVs
df = map_dfr(list.files(), read.csv, colClasses = "character")

```

---

<div class="post-metadata">

**Author:** ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)\
**Post date:** [June 2, 2020, 6:52pm UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/2 "2020-06-02T18:52:33Z")

</div>

I believe you can also use the column index with `colparsers`, e.g.

```julia
loadtable(path, colparsers=Dict(i => String for i in 1:n_cols))

```

---

<div class="post-metadata">

**Author:** ![anon64288406](https://avatars.discourse-cdn.com/v4/letter/a/da6949/32.png) [@anon64288406](https://discourse.julialang.org/u/anon64288406)\
**Post date:** [June 3, 2020, 1:10am UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/3 "2020-06-03T01:10:05Z")

</div>

This gives me:

> UndefVarError: n\_cols not defined

Where is `n_cols` supposed to be defined here?

---

<div class="post-metadata">

**Author:** ![joshday](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/joshday/32/368_2.png) [@joshday](https://discourse.julialang.org/u/joshday)\
**Post date:** [June 3, 2020, 11:17am UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/4 "2020-06-03T11:17:39Z")

</div>

`path` and `n_cols` are for you to define.

---

<div class="post-metadata">

**Author:** ![anon64288406](https://avatars.discourse-cdn.com/v4/letter/a/da6949/32.png) [@anon64288406](https://discourse.julialang.org/u/anon64288406)\
**Post date:** [July 6, 2020, 12:49pm UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/5 "2020-07-06T12:49:14Z")

</div>

But the number of columns varies by CSV, so me manually typing out and passing a Vector of Ints to `n_cols` would be tedious and error-prone if there are, say, 20 CSVs.

---

<div class="post-metadata">

**Author:** ![bernhard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bernhard/32/2619_2.png) [@bernhard](https://discourse.julialang.org/u/bernhard)\
**Post date:** [July 6, 2020, 2:51pm UTC](https://discourse.julialang.org/t/get-juliadb-loadtable-to-parse-all-columns-in-csvs-as-string/40620/6 "2020-07-06T14:51:17Z")

</div>

You can do `readline(my_csv_file)` to read the header. Then you can count how many times the delimiter occurs in the header. Generally this should be quite simple (except if the delimiter would occur in the header names, or if the first row of the file is not the header)
