# Bug with DelimtedFiles.readdlm when header=true

**URL:** <https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057>\
**Category:** General Usage\
**Tags:** package, csv\
**Created:** [February 26, 2021, 12:08am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057 "2021-02-26T00:08:48Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 12:08am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/1 "2021-02-26T00:08:48Z")

</div>

Hallo,

Thanks for providing an easy and fast way to read delimited files.

When I use _header=true_ as an option. In the output _Data_ it skips 2 columns instead of 1 which is a bug.

`Data, Header= DelimtedFiles.readdlm(path,',',header=true)`

I added an issue #39831

Many thanks for providing a fix.  
Joseph

---

<div class="post-metadata">

**Author:** ![stillyslalom](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stillyslalom/32/45687_2.png) [@stillyslalom](https://discourse.julialang.org/u/stillyslalom)\
**Post date:** [February 26, 2021, 12:17am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/2 "2021-02-26T00:17:08Z")

</div>

c.f. [header=true, skips 2 lines · Issue #39831 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/39831)

As I mentioned there, it’d really help if you could provide a small example CSV file that reproduces this behavior.

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 7:42am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/3 "2021-02-26T07:42:38Z")

</div>

Thanks Stilly for looking at the problem:

The input file which causes problem is attached  
[CSVfile](https://github.com/joseph-pollacco/SoilWaterToolbox/blob/master/INPUT/DataSoilHydraulic/Smap/Smap_Id_Select.csv)  
**The input file has the following format:**

> | Id | , SELECT\_1 |
> | --- | --- |
> | 11122911000031011, | 1 |
> | 11122911000031012, | 1 |
> | … | |

**The code**

```julia
Data, Header= DelimtedFiles.readdlm(path, ',' , header=true)

 println(Data[1,:]) = 11122911000031012

```

Which is not what is expected.

Hope that helps to debug,  
Joseph

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [February 26, 2021, 7:55am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/4 "2021-02-26T07:55:33Z")

</div>

I can see that a problem might arise if the data is treated as floating point. I am able to read the file in as integers:

```julia
julia> Data, Header= DelimitedFiles.readdlm("Smap_Id_Select.csv", ',' , Int, header=true);

julia> println(Int.(Data[1,:]))
[11122911000031011, 1]

```

whereas as floating-point:

```julia
julia> Data, Header= DelimitedFiles.readdlm("Smap_Id_Select.csv", ',' , header=true);

julia> println(Int.(Data[1,:]))
[11122911000031012, 1]

```

Won’t say that this is a `DelimitedFiles` bug exactly, rather a floating-point artifact

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 8:01am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/5 "2021-02-26T08:01:23Z")

</div>

Thanks for investigating. In this case I have only _Int64_, but what happens if **Id** is an integer and the data is _FLoat64_?

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 8:10am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/6 "2021-02-26T08:10:35Z")

</div>

My work around is to avoid using Header as follow:

```julia
         # Read data
            Data = DelimitedFiles.readdlm("Smap_Id_Select.csv", ',')

         # Read header
            Header = Data[1,1:end]

         # Remove first row
            Data = Data[1:end.≠1,1:end]

```

---

<div class="post-metadata">

**Author:** ![jishnub](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jishnub/32/33620_2.png) [@jishnub](https://discourse.julialang.org/u/jishnub)\
**Post date:** [February 26, 2021, 8:16am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/7 "2021-02-26T08:16:08Z")

</div>

Such rounding might be happening at more places than the first row, so removing the first row is not the solution. If you have data of mixed types then it’s best to read it in by specifying the element type to be `Any`.

```julia
julia> Data, Header= DelimitedFiles.readdlm("Smap_Id_Select.csv", ',' , Any, header=true);

julia> Data[1:2, :]
2×2 Array{Any,2}:
 11122911000031011 1
 11122911000031012 1

julia> Data[1,1] |> typeof
Int64

```

For more sophisticated cases you may consider using a `DataFrame`

```julia
julia> using CSV, DataFrames

julia> df = DataFrame(CSV.File("Smap_Id_Select.csv"))
16281×2 DataFrame
   Row │ Id SELECT_1 
       │ Int64 Int64    
───────┼─────────────────────────────
     1 │ 11122911000031011 1
     2 │ 11122911000031012 1
     3 │ 11122911000031013 1
     4 │ 11122911000031014 1
[...]

```

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 8:19am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/8 "2021-02-26T08:19:40Z")

</div>

**Requiring some further explanations.**  
When I remove the _first row_ I am removing the row which has the **Headings**.

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [February 26, 2021, 8:27am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/9 "2021-02-26T08:27:45Z")

</div>

What is the fastest way to _read .csv_ =\> **Array** and not DataFrames?

---

<div class="post-metadata">

**Author:** ![greg\_plowman](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/greg_plowman/32/8100_2.png) [@greg\_plowman](https://discourse.julialang.org/u/greg_plowman)\
**Post date:** [February 26, 2021, 8:37am UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/10 "2021-02-26T08:37:23Z")

</div>

The issue with specifying `header=true` but no type is that by default, type is assumed to be `Float64` and you get rounding errors on column 1 data.

In your example this can be fixed by specifying `header=true` and type `Int` for data.

But …

> In this case I have only Int64, but what happens if Id is an integer and the data is FLoat64?

Well, here you can specify `header=true` with type `Any`.

Or use `CSV`, which can do a better job of inferring column types, or you can explicitly specify them.

Your solution of reading without `header=true` and manually slicing out the first row works because the type is automatically inferred as `Any` since the first row (header) contains strings and the data rows are numeric.

BTW, I think you can simplify your slicing:

```julia
# Read header
    Header = Data[1,1:end]

 # Remove first row
    Data = Data[1:end.≠1,1:end]

```

```julia
Header = Data[1, :]
Data = Data[2:end, :]

```

> What is the fastest way to _read .csv_ =\> **Array** and not DataFrames?

Checkout `CSV`:

> **[GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other...](https://github.com/JuliaData/CSV.jl)**
>
> Utility library for working with CSV and other delimited files in the Julia programming language - GitHub - JuliaData/CSV.jl: Utility library for working with CSV and other delimited files in the J...

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [March 4, 2021, 9:43pm UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/11 "2021-03-04T21:43:32Z")

</div>

I would like to thank you for providing us such a **DelimtedFiles** which is in the core Julia package and therefore I expect that the **DelimtedFiles** works as expected with no surprises.

To my understanding in the documentation  
[DelimitedFiles](https://docs.julialang.org/en/v1/stdlib/DelimitedFiles/#DelimitedFiles.readdlm-Tuple%7BAny,AbstractChar,Type,AbstractChar%7D)

There is no where written that one must putting _Any_ or else you might not get the expected results.

From my perspective **Delimited Files** has a **bug** which needs to be fixed so that users can get the results as the documentation promises.

Once again thanks for providing a free tool,  
Joseph

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [March 4, 2021, 9:45pm UTC](https://discourse.julialang.org/t/bug-with-delimtedfiles-readdlm-when-header-true/56057/12 "2021-03-04T21:45:49Z")

</div>

My comment is that if **DelimitedFiles** is outdated it may be a good idea to replace it with **CSV.jl.** I understand that there must be some build tools to easily convert DataFrameworks into Array.
