# Read array of strings into Dictionary of DataFrames

**URL:** <https://discourse.julialang.org/t/read-array-of-strings-into-dictionary-of-dataframes/43202>\
**Category:** Performance\
**Tags:** question, dataframes\
**Created:** [July 17, 2020, 12:54am UTC](https://discourse.julialang.org/t/read-array-of-strings-into-dictionary-of-dataframes/43202 "2020-07-17T00:54:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![evad](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/evad/32/15756_2.png) [@evad](https://discourse.julialang.org/u/evad)\
**Post date:** [July 17, 2020, 12:54am UTC](https://discourse.julialang.org/t/read-array-of-strings-into-dictionary-of-dataframes/43202/1 "2020-07-17T00:54:00Z")

</div>

Hi,  
I have an array of strings where each string is a stream of multiple tables. And the final step of my function is to create and return a dictionary:

```julia
function readTables(str)		
		nStrs = split(str,"\r\n%T\t")
		fHead=nStrs[1]
		fTail="%E"
		nStrs[end]=replace(nStrs[end],r"%E.*"=>"")
		tableStrings=nStrs[2:end]
		tableNames=match.(r"^[A-Z]+",tableStrings)
		tableNames=[x.match for x in tableNames]
		dictTables=Dict(tableNames.=>CSV.read.(IOBuffer.(tableStrings),delim="\t",header=2, type=String))
	end
	#map(readTables,values(contentP6))

```

But when i run the map line above, the browser just hangs. I’m using Julia 1.4.2 / Pluto via Chrome. The size of the string array `values(contentP6)` is about 1.7Gb, 461 elements.

May I know:

1. Is Dict of 2D Arrays the right way to store these tables?
2. How to improve the code for better performance?

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [July 17, 2020, 6:19am UTC](https://discourse.julialang.org/t/read-array-of-strings-into-dictionary-of-dataframes/43202/2 "2020-07-17T06:19:19Z")

</div>

1. That’s up to you — it could be a sensible solution.

2. I would consider some solution that does not need to read the file and then make chunks — eg `Mmap` the file, find the delimiters, pass a view to `CSV.File`, and read that.

In any case, I would factor this out to smaller functions so that I could benchmark each piece.
