# CSV won't read tab separated file

**URL:** <https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120>\
**Category:** General Usage\
**Tags:** csv\
**Created:** [March 4, 2024, 9:21am UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120 "2024-03-04T09:21:15Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 9:21am UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/1 "2024-03-04T09:21:15Z")

</div>

I have a large text file that describes itself as

```julia
--
-- PostgreSQL database dump
--

-- Dumped from database version 12.16
-- Dumped by pg_dump version 12.16

```

It contains several blocks of tab-separated data which represent the data from each database table. I’ve copied one of these blocks (using the VSCode editor) into a stand-alone file. This contains 15 tab-separated columns and 761,065 rows. First row is a header. I cannot coax CSV into reading this file.

```julia
using CSV, DataFrames, StringEncodings

entries = ["id", "identifier", "title", "description", "currency", "amount_awarded", "award_date", "last_modified", "funding_org_id", "good_cause_area_id", "local_authority_id", "recipient_org_id", "region_id", "uk_constituency_id", "ward_id"]
    # Names of DF columns

df1 = DataFrame([name => String[] for name in entries])
for line in eachline("newex.csv")
    list = findall(r"\t", line)
    l = length(list)
    if l !== 14 # Check that all lines have exactly 14 tab separators
        println(l)
    end
    row = String[]
    append!(row, [chop(line[1:first(list[1])], head=0, tail=1)], [chop(line[first(list[x-1]):first(list[x])], head=1, tail=1) for x = 2:l], [chop(line[first(list[l]):end], head=1, tail=0)])
    push!(df1, row)
end
deleteat!(df1, 1) # first row of datafile contains headers.
println(describe(df1))

#df2 = CSV.File(open("newex.csv", enc"ISO-8859-1")) |> DataFrame
df2 = CSV.read("newex.csv", DataFrame)
println(describe(df2))

```

My “brute force and ignorance” method produces a viable dataframe, but CSV.read fails. I’ve tried specifying `delim="\t"` and have experimented with StringEncodings but nothing I’ve tried makes any difference.  
Obviously, I’ve managed to read the datafile, so this post is out of curiosity more than anything.

Here is the output of running the above script.

```julia
15×7 DataFrame
 Row │ variable mean min median max nmissing eltype   
     │ Symbol Nothing String Nothing String Int64 DataType 
─────┼────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
   1 │ id 100 999999 0 String
   2 │ identifier 000406 YH201117526 0 String
   3 │ title " " (working title) ☑ Other 0 String
   4 │ description ! (Hollywood Blues) \uf0fc Musicians, Dancers, S… 0 String
   5 │ currency GBP GBP 0 String
   6 │ amount_awarded 0 999999 0 String
   7 │ award_date 1995-01-06 2023-12-20 0 String
   8 │ last_modified 2021-12-09 12:53:26.999321+00 2024-01-11 13:07:00.623853+00 0 String
   9 │ funding_org_id 1 9 0 String
  10 │ good_cause_area_id 1 8 0 String
  11 │ local_authority_id 1 \\N 0 String
  12 │ recipient_org_id 1 99999 0 String
  13 │ region_id 1 \\N 0 String
  14 │ uk_constituency_id 1 \\N 0 String
  15 │ ward_id 

```

Then I get literally thousands of warnings from CSV as it tries to read each line of the file:

```julia
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548111. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548112. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548113. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548114. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548115. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548116. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548117. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 3 warning: only found 15 / 16 columns around data row: 548118. Filling remaining columns with `missing`
└ @ CSV C:\U
...
┌ Warning: thread = 3: too many warnings, silencing any further warnings
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:591
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 195709. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 195709. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 195710. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 195710. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 195711. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 195711. Filling remaining columns with `missing`
...
┌ Warning: thread = 1: too many warnings, silencing any further warnings
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:591
┌ Warning: thread = 1: too many warnings, silencing any further warnings
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:591
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601498. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601499. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601500. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601501. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601502. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 15 / 16 columns around data row: 601503. Filling remaining columns with `missing`
...
┌ Warning: thread = 1: too many warnings, silencing any further warnings
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:591
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601514. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601515. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601516. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601517. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601518. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601519. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601520. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601521. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
...
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601595. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601596. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601597. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 15 / 16 columns around data row: 601598. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4: too many warnings, silencing any further warnings
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:591

```

And then the crash.

```julia
ERROR: LoadError: UndefRefError: access to undefined reference
Stacktrace:
  [1] getindex
    @ .\essentials.jl:13 [inlined]
  [2] iterate
    @ .\array.jl:945 [inlined]
  [3] _foldl_impl
    @ .\reduce.jl:60 [inlined]
  [4] foldl_impl
    @ .\reduce.jl:48 [inlined]
  [5] mapfoldl_impl
    @ .\reduce.jl:44 [inlined]
  [6] _mapreduce_dim
    @ .\reducedim.jl:362 [inlined]
  [7] mapreduce
    @ .\reducedim.jl:357 [inlined]
  [8] _sum
    @ .\reducedim.jl:1015 [inlined]
  [9] sum
    @ .\reducedim.jl:1011 [inlined]
 [10] _simple_count
    @ .\reduce.jl:1357 [inlined]
 [11] _count
    @ .\reducedim.jl:442 [inlined]
 [12] count
    @ .\reducedim.jl:440 [inlined]
 [13] #43
    @ .\none:0 [inlined]
 [14] MappingRF
    @ .\reduce.jl:100 [inlined]
 [15] _foldl_impl(op::Base.MappingRF{SentinelArrays.var"#43#44"{typeof(ismissing)}, Base.BottomRF{typeof(Base.add_sum)}}, init::Base._InitialValue, itr::Vector{Vector{String}})      
    @ Base .\reduce.jl:58
 [16] foldl_impl
    @ .\reduce.jl:48 [inlined]
 [17] mapfoldl_impl
    @ .\reduce.jl:44 [inlined]
 [18] mapfoldl
    @ .\reduce.jl:175 [inlined]
 [19] mapreduce
    @ .\reduce.jl:307 [inlined]
 [20] sum
    @ .\reduce.jl:535 [inlined]
 [21] sum
    @ .\reduce.jl:564 [inlined]
 [22] count
    @ C:\Users\TGebbels\.julia\packages\SentinelArrays\1kRo4\src\chainedvector.jl:787 [inlined]
 [23] (::DataFrames.var"#76#84"{Vector{Any}})(col::SentinelArrays.ChainedVector{String, Vector{String}})
    @ DataFrames C:\Users\TGebbels\.julia\packages\DataFrames\58MUJ\src\abstractdataframe\abstractdataframe.jl:723
 [24] iterate
    @ .\generator.jl:47 [inlined]
 [25] collect_to!(dest::Vector{Dict{Symbol, Any}}, itr::Base.Generator{DataFrames.DataFrameColumns{DataFrame}, DataFrames.var"#76#84"{Vector{Any}}}, offs::Int64, st::Int64)
    @ Base .\array.jl:892
 [26] collect_to_with_first!
    @ .\array.jl:870 [inlined]
 [27] collect(itr::Base.Generator{DataFrames.DataFrameColumns{DataFrame}, DataFrames.var"#76#84"{Vector{Any}}})
    @ Base .\array.jl:844
 [28] map
    @ .\abstractarray.jl:3313 [inlined]
 [29] _describe(df::DataFrame, stats::Vector{Any})
    @ DataFrames C:\Users\TGebbels\.julia\packages\DataFrames\58MUJ\src\abstractdataframe\abstractdataframe.jl:712
 [30] describe(df::DataFrame; cols::Function)
    @ DataFrames C:\Users\TGebbels\.julia\packages\DataFrames\58MUJ\src\abstractdataframe\abstractdataframe.jl:668
 [31] top-level scope
    @ c:\Users\TGebbels\OneDrive - xxxxxxxx\Documents\DCMS Database\TestCSVread.jl:37

```

I don’t have a mechanism to share the datafile but I would if I could.  
Here are the first few lines:

```julia
id	identifier	title	description	currency	amount_awarded	award_date	last_modified	funding_org_id	good_cause_area_id	local_authority_id	recipient_org_id	region_id	uk_constituency_id	ward_id
99402	DCMS-tnlcomfund-0023093253	Activity Adventure Trai	The organisation will use the funding to purchase and construct an Activity Adventure Trail. This will enable the organisation to provide a safe outdoor recreational facility which the area lacks currently to tackle the identified risk of isolation and anti social behaviour amongst the local children.	GBP	0	2015-07-30	2022-06-07 13:24:35.117395+00	5	8	306	63764	3	317	6032
99403	DCMS-tnlcomfund-0023093271	The SPHERE Project	The organisation will use the funding to run a series of activity sessions focussing on physical mental and sexual health. This will enable the organisation to tackle the ongoing health and substance misuse issues that arise as a result of deprivation in the area.	GBP	0	2015-07-30	2022-06-07 13:24:35.129602+00	5	8	76	46015	3	91	11
101347	DCMS-tnlcomfund-0023093273	Crookston Environmental and Outdoor Project	Crookston Environmental and Outdoor Project	GBP	0	2015-08-12	2022-06-07 13:24:35.135671+00	5	8	205	64760	1	545	3978
97860	DCMS-tnlcomfund-0023093278	A4A - Paulsgrove and Wymering Connect	The organisation will use the funding to work with over 40 local agencies to create an online directory giving residents information about local activities and services. This will enable the group to provide easy access to all local residents with updated programmes of activities and information about accessing these services to improve wellbeing and community engagement in the area.	GBP	0	2015-07-16	2022-06-07 13:24:35.141859+00	5	8	141	62886	9	391	767
112863	DCMS-tnlcomfund-0023207672	The Engage Programme	The project will provide practical life skills and training to teenage boys who lack positive male role models in their lives. This will improve the young people’s self-esteem resilience and motivation.	GBP	0	2017-10-12	2022-06-07 15:10:41.356337+00	5	8	366	68266	7	602	5357

```

---

<div class="post-metadata">

**Author:** ![Yuan-Ru-Lin](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yuan-ru-lin/32/46068_2.png) [@Yuan-Ru-Lin](https://discourse.julialang.org/u/Yuan-Ru-Lin)\
**Post date:** [March 4, 2024, 10:01am UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/2 "2024-03-04T10:01:40Z")

</div>

I ran the script with `newex.csv` being the first few lines you provided.

The result is

```julia
15×7 DataFrame
 Row │ variable mean min median max nmissing eltype
     │ Symbol Union… Any Any Any Int64 DataType
─────┼────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
   1 │ id 102175.0 97860 99403.0 112863 0 Int64
   2 │ identifier DCMS-tnlcomfund-0023093253 DCMS-tnlcomfund-0023207672 0 String31
   3 │ title A4A - Paulsgrove and Wymering Co… The SPHERE Project 0 String
   4 │ description Crookston Environmental and Outd… The project will provide practic… 0 String
   5 │ currency GBP GBP 0 String3
   6 │ amount_awarded 0.0 0 0.0 0 0 Int64
   7 │ award_date 2015-07-16 2015-07-30 2017-10-12 0 Date
   8 │ last_modified 2022-06-07 13:24:35.117395+00 2022-06-07 15:10:41.356337+00 0 String31
   9 │ funding_org_id 5.0 5 5.0 5 0 Int64
  10 │ good_cause_area_id 8.0 8 8.0 8 0 Int64
  11 │ local_authority_id 218.8 76 205.0 366 0 Int64
  12 │ recipient_org_id 61138.2 46015 63764.0 68266 0 Int64
  13 │ region_id 4.6 1 3.0 9 0 Int64
  14 │ uk_constituency_id 389.2 91 391.0 602 0 Int64
  15 │ ward_id 3229.0 11 3978.0 6032 0 Int64

```

Maybe the error was due to corrupted data in the rest of the file?

* * *

Here is the result of `versioninfo()`

```julia
Julia Version 1.10.1
Commit 7790d6f0641 (2024-02-13 20:41 UTC)
Build Info:
  Official https://julialang.org/ release
Platform Info:
  OS: macOS (arm64-apple-darwin22.4.0)
  CPU: 10 × Apple M2 Pro
  WORD_SIZE: 64
  LIBM: libopenlibm
  LLVM: libLLVM-15.0.7 (ORCJIT, apple-m1)
Threads: 1 default, 0 interactive, 1 GC (on 6 virtual cores)

```

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 11:29am UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/3 "2024-03-04T11:29:09Z")

</div>

My simplistic approach can successfully read the whole file.  
Some of these fields contain “\\N” to indicate missing values. (eg Local Authority, Region and Constituency).

I process these lines like this:

```julia
    df3.local_authority_id = [row.local_authority_id == "\\N" ? missing : parse(Int64, row.local_authority_id) for row in eachrow(df1)]
    df3.region_id = [row.region_id == "\\N" ? missing : parse(Int64, row.region_id) for row in eachrow(df1)]
    df3.uk_constituency_id = [row.uk_constituency_id == "\\N" ? missing : parse(Int64, row.uk_constituency_id) for row in eachrow(df1)]
    df3.ward_id = [row.ward_id == "\\N" ? missing : parse(Int64, row.ward_id) for row in eachrow(df1)]

```

Again, perhaps, not elegant, but effective.

Might it be these “\\N” values that are confusing CSV?

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [March 4, 2024, 12:32pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/4 "2024-03-04T12:32:21Z")

</div>

That wouldn’t surprise me, have you tried telling `CSV` about them?

```julia
    • missingstring: either a nothing, String, or Vector{String} to use as sentinel values that will be parsed as missing; if nothing is passed, no sentinel/missing values will be parsed; by default,
       missingstring="", which means only an empty field (two consecutive delimiters) is considered missing

```

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 12:44pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/5 "2024-03-04T12:44:23Z")

</div>

So, I tried

```julia
df2 = CSV.read("newex.csv", missingstring="\\N", DataFrame)

```

and

```julia
df2 = CSV.read("newex.csv", DataFrame; missingstring=["\\N", ""])

```

but still got the same warnings and errors as above.

If I try

```julia
df2 = CSV.read("newex.csv", missingstring="\N", DataFrame)

```

as suggested [here](https://csv.juliadata.org/v0.2.4/), I get

```julia
df2 = CSV.read("newex.csv", missingstring="\N", DataFrame)
# └┘ ── invalid escape sequence
Stacktrace:
 [1] top-level scope
   @ c:\Users\TGebbels\OneDrive - xxxxxxxxxxx\Documents\DCMS Database\TestCSVread.jl:36
in expression starting at c:\Users\TGebbels\OneDrive - xxxxxxxxxx\Documents\DCMS Database\TestCSVread.jl:36

```

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [March 4, 2024, 12:48pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/6 "2024-03-04T12:48:18Z")

</div>

If there are actually two backslashes in the raw data you need `missingstring="\\\\N"`

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 12:50pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/7 "2024-03-04T12:50:28Z")

</div>

I don’t think so, because my method above worked with `"\\N"`.

Just to be sure, I tried

```julia
df2 = CSV.read("newex.csv", DataFrame; missingstring=["\\\\N", ""])

```

but the results were the same.

---

<div class="post-metadata">

**Author:** ![nilshg](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/nilshg/32/2283_2.png) [@nilshg](https://discourse.julialang.org/u/nilshg)\
**Post date:** [March 4, 2024, 1:01pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/8 "2024-03-04T13:01:12Z")

</div>

> [@TimG](#):
>
> I don’t think so, because my method above worked with `"\\N"`.

I edited your file as follows:

![image](https://global.discourse-cdn.com/julialang/original/3X/d/2/d27d518a4293d38ad82aace4fdf8c888ecbecbd0.png)

Then:

```julia
julia> CSV.read(f, DataFrame)
5×15 DataFrame
 Row │ id identifier title description currency amount_awarded award_date last_modified funding_org_id good_cause_ ⋯
     │ Int64 String31 String String String3 Int64 Dates.Date String31 Int64 Int64 ⋯
─────┼─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
   1 │ 99402 DCMS-tnlcomfund-0023093253 Activity Adventure Trai The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.117395+00 5 ⋯
   2 │ 99403 DCMS-tnlcomfund-0023093271 \\\\N The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.129602+00 5
   3 │ 101347 DCMS-tnlcomfund-0023093273 Crookston Environmental and Outd… Crookston Environmental and Outd… GBP 0 2015-08-12 2022-06-07 13:24:35.135671+00 5
   4 │ 97860 DCMS-tnlcomfund-0023093278 A4A - Paulsgrove and Wymering Co… The organisation will use the fu… GBP 0 2015-07-16 2022-06-07 13:24:35.141859+00 5
   5 │ 112863 DCMS-tnlcomfund-0023207672 The Engage Programme The project will provide practic… GBP 0 2017-10-12 2022-06-07 15:10:41.356337+00 5 ⋯
                                                                                                                                                                                                  6 columns omitted

julia> CSV.read(f, DataFrame; missingstring = "\\N")
5×15 DataFrame
 Row │ id identifier title description currency amount_awarded award_date last_modified funding_org_id good_cause_ ⋯
     │ Int64 String31 String String String3 Int64 Dates.Date String31 Int64 Int64 ⋯
─────┼─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
   1 │ 99402 DCMS-tnlcomfund-0023093253 Activity Adventure Trai The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.117395+00 5 ⋯
   2 │ 99403 DCMS-tnlcomfund-0023093271 \\\\N The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.129602+00 5
   3 │ 101347 DCMS-tnlcomfund-0023093273 Crookston Environmental and Outd… Crookston Environmental and Outd… GBP 0 2015-08-12 2022-06-07 13:24:35.135671+00 5
   4 │ 97860 DCMS-tnlcomfund-0023093278 A4A - Paulsgrove and Wymering Co… The organisation will use the fu… GBP 0 2015-07-16 2022-06-07 13:24:35.141859+00 5
   5 │ 112863 DCMS-tnlcomfund-0023207672 The Engage Programme The project will provide practic… GBP 0 2017-10-12 2022-06-07 15:10:41.356337+00 5 ⋯
                                                                                                                                                                                                  6 columns omitted

julia> CSV.read(f, DataFrame; missingstring = "\\\\N")
5×15 DataFrame
 Row │ id identifier title description currency amount_awarded award_date last_modified funding_org_id good_cause_ ⋯
     │ Int64 String31 String? String String3 Int64 Dates.Date String31 Int64 Int64 ⋯
─────┼─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
   1 │ 99402 DCMS-tnlcomfund-0023093253 Activity Adventure Trai The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.117395+00 5 ⋯
   2 │ 99403 DCMS-tnlcomfund-0023093271 missing The organisation will use the fu… GBP 0 2015-07-30 2022-06-07 13:24:35.129602+00 5
   3 │ 101347 DCMS-tnlcomfund-0023093273 Crookston Environmental and Outd… Crookston Environmental and Outd… GBP 0 2015-08-12 2022-06-07 13:24:35.135671+00 5
   4 │ 97860 DCMS-tnlcomfund-0023093278 A4A - Paulsgrove and Wymering Co… The organisation will use the fu… GBP 0 2015-07-16 2022-06-07 13:24:35.141859+00 5
   5 │ 112863 DCMS-tnlcomfund-0023207672 The Engage Programme The project will provide practic… GBP 0 2017-10-12 2022-06-07 15:10:41.356337+00 5 ⋯
                                                                                                                                                                                                  6 columns omitted

```

Unless you want to share the whole file there’s little we can do to help you. Given the first few lines of your file work you should probably just read the first `n` lines, increasing `n` until you hit the error to work out where your file is corrupted.

---

<div class="post-metadata">

**Author:** ![johnh](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnh/32/3615_2.png) [@johnh](https://discourse.julialang.org/u/johnh)\
**Post date:** [March 4, 2024, 1:59pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/9 "2024-03-04T13:59:53Z")

</div>

$DEITY help me for saying this. pg\_dump creates a file that can be read by another Postgres database. Could you create a local Postgres database and read this file in?  
Then us Julia to extract the values from the database?

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 3:00pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/10 "2024-03-04T15:00:34Z")

</div>

I don’t have Postgres. The file was produced by a 3rd party.  
I have read the file. I’m just wondering why CSV can’t.

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 3:01pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/11 "2024-03-04T15:01:59Z")

</div>

> [@nilshg](#):
>
> Unless you want to share the whole file there’s little we can do to help you. Given the first few lines of your file work you should probably just read the first `n` lines, increasing `n` until you hit the error to work out where your file is corrupted.

I’d like to share the file but am unsure of a mechanism.  
Tim

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 3:08pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/12 "2024-03-04T15:08:49Z")

</div>

> [@nilshg](#):
>
> Given the first few lines of your file work you should probably just read the first `n` lines, increasing `n` until you hit the error to work out where your file is corrupted.

It looks like line 334 is where the problem first arises:

```julia
println("limit=333")
CSV.read("newex.csv", DataFrame; limit=333)#; missingstring="\\\\N")
println("limit=334")
CSV.read("newex.csv", DataFrame; limit=334)#; missingstring="\\\\N")

```

This produces:

```julia
limit=333
limit=334
┌ Warning: thread = 4 warning: only found 8 / 15 columns around data row: 496. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 8 / 15 columns around data row: 515. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 4 warning: only found 8 / 15 columns around data row: 515. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
┌ Warning: thread = 1 warning: only found 8 / 15 columns around data row: 515. Filling remaining columns with `missing`
└ @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:586
ERROR: LoadError: TaskFailedException

    nested task error: BoundsError: attempt to access 334-element Vector{UInt32} at index [335]
    Stacktrace:
     [1] setindex!
       @ .\array.jl:1021 [inlined]
     [2] checkpooled!(::Type{Union{Missing, String3}}, pertaskcolumns::Vector{Vector{CSV.Column}}, col::CSV.Column, j::Int64, ntasks::Int64, nrows::Int64, ctx::CSV.Context)
       @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:514
     [3] multithreadpostparse(ctx::CSV.Context, ntasks::Int64, pertaskcolumns::Vector{Vector{CSV.Column}}, rows::Vector{Int64}, finalrows::Int64, j::Int64, col::CSV.Column)
       @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:433
     [4] (::CSV.var"#35#40"{CSV.Context, Int64, Vector{Vector{CSV.Column}}, Vector{Int64}, Int64, Int64, CSV.Column})()
       @ CSV C:\Users\TGebbels\.julia\packages\WorkerUtilities\ey0fP\src\WorkerUtilities.jl:384

...and 2 more exceptions.

Stacktrace:
 [1] sync_end(c::Channel{Any})
   @ Base .\task.jl:448
 [2] macro expansion
   @ .\task.jl:480 [inlined]
 [3] CSV.File(ctx::CSV.Context, chunking::Bool)
   @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:282
 [4] File
   @ C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:227 [inlined]
 [5] #File#32
   @ C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\file.jl:223 [inlined]
 [6] read(source::String, sink::Type; copycols::Bool, kwargs::@Kwargs{limit::Int64})
   @ CSV C:\Users\TGebbels\.julia\packages\CSV\aoJqo\src\CSV.jl:117
 [7] top-level scope
   @ c:\Users\TGebbels\OneDrive - xxxxxx\Documents\DCMS Database\TestCSVread.jl:39
in expression starting at c:\Users\TGebbels\OneDrive - xxxxxx\Documents\DCMS Database\TestCSVread.jl:39

```

This is line 334 from the datafile:

```julia
98021	DCMS-tnlcomfund-0023093356	Community Project	The group will use the funding to make improvements to the flooring and purchase of tables and chairs for the frequently used meeting room. This will improve community buildings and increase usage of facilities.	GBP	0	2015-07-23	2022-06-07 13:24:35.23891+00	5	8	176	62912	8	235	137

```

```julia

```

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [March 4, 2024, 3:36pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/13 "2024-03-04T15:36:34Z")

</div>

Also share line 333, which could trigger the problem.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [March 4, 2024, 4:11pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/14 "2024-03-04T16:11:12Z")

</div>

> [@TimG](#):
>
> share the file

Upload the file to gitub, dropbox, google drive, or whatever service you like and then post a link here

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [March 4, 2024, 4:15pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/15 "2024-03-04T16:15:22Z")

</div>

> [@TimG](#):
>
> `Warning: thread = 1 warning: only found 8 / 15 columns around data row: 515`

I think it is unlikely that the problem is with line 334 because it has progressed to at least line 515.

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 4:15pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/16 "2024-03-04T16:15:56Z")

</div>

Line 333

```julia
505	360G-ACE-01-L090504154L-13	Grant to Ambit Magazine	Ambit Tour of the South West	GBP	4951	2005-03-01	2022-04-13 13:56:52.789849+00	1	1	45	35	2	56	305

```

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 4:19pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/17 "2024-03-04T16:19:21Z")

</div>

Github says:

```julia
Yowza, that’s a big file. Try again with a file smaller than 25MB.

```

It’s 244MB.

---

<div class="post-metadata">

**Author:** ![Jeff\_Emanuel](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jeff_emanuel/32/15440_2.png) [@Jeff\_Emanuel](https://discourse.julialang.org/u/Jeff_Emanuel)\
**Post date:** [March 4, 2024, 4:21pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/18 "2024-03-04T16:21:07Z")

</div>

[https://www.google.com/search?q=large+file+sharing+service](https://www.google.com/search?q=large+file+sharing+service)

---

<div class="post-metadata">

**Author:** ![Dan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dan/32/42581_2.png) [@Dan](https://discourse.julialang.org/u/Dan)\
**Post date:** [March 4, 2024, 4:27pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/19 "2024-03-04T16:27:16Z")

</div>

And add `ntasks=1` to `CSV.read` to avoid multitasking. This would allow pinpointing the offending line.

```julia
CSV.read("newex.csv", DataFrame; limit=333, ntasks=1)

```

---

<div class="post-metadata">

**Author:** ![TimG](https://avatars.discourse-cdn.com/v4/letter/t/82dd89/32.png) [@TimG](https://discourse.julialang.org/u/TimG)\
**Post date:** [March 4, 2024, 4:27pm UTC](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120/20 "2024-03-04T16:27:34Z")

</div>

Google wouldn’t let me upload it as a CSV file so I changed the file type to .txt - still the same file, though.

[newex.txt](https://drive.google.com/file/d/1b_bcGeEkwFKnabezNGhsSI2beyHhq9w2/view?usp=drive_link)

[Next page](https://discourse.julialang.org/t/csv-wont-read-tab-separated-file/111120.md?page=2)
