# How to parse text between specific characters in Julia?

**URL:** <https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284>\
**Category:** New to Julia\
**Created:** [August 28, 2023, 9:45am UTC](https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284 "2023-08-28T09:45:24Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Luigi\_Marongiu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/luigi_marongiu/32/7909_2.png) [@Luigi\_Marongiu](https://discourse.julialang.org/u/Luigi_Marongiu)\
**Post date:** [August 28, 2023, 9:45am UTC](https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284/1 "2023-08-28T09:45:24Z")

</div>

Hello,  
I would like to create a text files (a plain text txt or a tab-separated file tsv) with modular information. I have been thinking of including the information within some markers, for instance `*=` and `=*` so that the text within these markers works as a unit and multiple units can be placed in the file. Somenting like this:

```julia
*=
NAME_Abel
DOB_03-09-1960
*=
*=
NAME_Bernard
DOB_12-12-1972
*=

```

Is it possible to load these chunks of text one group at the time in Julia? The idea is to read the file, parse in the first group `*= ... =*`, then parse the other and so forth.  
Thank you

---

<div class="post-metadata">

**Author:** ![oheil](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oheil/32/220745_2.png) [@oheil](https://discourse.julialang.org/u/oheil)\
**Post date:** [August 28, 2023, 10:18am UTC](https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284/2 "2023-08-28T10:18:51Z")

</div>

How about using XML for your task?  
Example:

```julia
julia> using LightXML

julia> xdoc = parse_file("xml_example.txt");

julia> xroot = root(xdoc);

julia> for ele in get_elements_by_tagname(xroot, "group")
       println(content(ele))
       end

NAME_Abel
DOB_03-09-1960

NAME_Bernard
DOB_12-12-1972

```

The file `xml_example.txt` looks like:

```julia
<?xml version="1.0" encoding="UTF-8"?>
<groups>
<group>
NAME_Abel
DOB_03-09-1960
</group>
<group>
NAME_Bernard
DOB_12-12-1972
</group>
</groups>

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [August 28, 2023, 12:49pm UTC](https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284/3 "2023-08-28T12:49:35Z")

</div>

> [@Luigi\_Marongiu](#):
>
> I would like to create a text files (a plain text txt or a tab-separated file tsv) with modular information.

Although you could definitely write a parser in Julia for any format you want, like @oheil I would encourage you to use a standard format, via one of Julia’s many existing parsers, rather than writing your own.

For structured data in a human-readable text-based format, a common choice is [JSON format](https://en.wikipedia.org/wiki/JSON) (using [JSON.jl](https://github.com/JuliaIO/JSON.jl)). XML is also an option, but in my opinion that’s overly complicated for relatively simple data structures.

There are also binary formats, like the BSON analogue of JSON (via [BSON.jl](https://github.com/JuliaIO/BSON.jl)), [Protocol Buffers](https://en.wikipedia.org/wiki/Protocol_Buffers) (using [ProtoBuf.jl](https://github.com/JuliaIO/ProtoBuf.jl)), the HDF5 format (via [HDF5.jl](https://github.com/JuliaIO/HDF5.jl)) which is more oriented towards large numeric arrays, and many others, and JLD format (which is Julia-oriented wrapper around HDF5, via [JLD.jl](https://github.com/JuliaIO/JLD.jl) or [JDL2.jl](https://github.com/JuliaIO/JLD2.jl)), among many others.

---

<div class="post-metadata">

**Author:** ![Luigi\_Marongiu](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/luigi_marongiu/32/7909_2.png) [@Luigi\_Marongiu](https://discourse.julialang.org/u/Luigi_Marongiu)\
**Post date:** [August 28, 2023, 3:01pm UTC](https://discourse.julialang.org/t/how-to-parse-text-between-specific-characters-in-julia/103284/4 "2023-08-28T15:01:16Z")

</div>

Thank you. I’ll look into them.
