# Parse string representation of a vector into floats

**URL:** <https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025>\
**Category:** General Usage\
**Tags:** question\
**Created:** [September 4, 2024, 9:43am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025 "2024-09-04T09:43:52Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 9:43am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/1 "2024-09-04T09:43:53Z")

</div>

How can I parse the following string:

```julia
"Float32[0.0, 8.455541, 16.61883, 24.443338, 31.901924, 38.97408, 45.642822, 46.02448, 47.268967, 47.10146, 47.10146]"

```

as vector of Float32 values?

---

<div class="post-metadata">

**Author:** ![fredrikekre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/fredrikekre/32/1688_2.png) [@fredrikekre](https://discourse.julialang.org/u/fredrikekre)\
**Post date:** [September 4, 2024, 9:49am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/2 "2024-09-04T09:49:10Z")

</div>

```julia
julia> str = "Float32[0.0, 8.455541]";

julia> m = match(r"\[(.*)\]", str)
RegexMatch("[0.0, 8.455541]", 1="0.0, 8.455541")

julia> strs = split(m[1], ','; keepempty=false)
2-element Vector{SubString{String}}:
 "0.0"
 " 8.455541"

julia> parse.(Float32, strs)
2-element Vector{Float32}:
 0.0
 8.455541

```

(Why do you have a string like that in the first place?)

---

<div class="post-metadata">

**Author:** ![screw\_dog](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/screw_dog/32/48119_2.png) [@screw\_dog](https://discourse.julialang.org/u/screw_dog)\
**Post date:** [September 4, 2024, 10:09am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/3 "2024-09-04T10:09:29Z")

</div>

Just to add: the highly _not_ _recommended_ way to do this is to use `Meta.parse` and `eval`, which allows you to treat a string as Julia code. But this is almost always not the best way to do things.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 11:28am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/4 "2024-09-04T11:28:33Z")

</div>

> [@fredrikekre](#):
>
> (Why do you have a string like that in the first place?)

I am importing a CSV file which itself was exported from an .arrow file.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 12:14pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/5 "2024-09-04T12:14:37Z")

</div>

This mostly works, but sometimes it does not. Example:

```julia
"Float32[0.0, 3.5440563f-28, -8.439492f-25, 7.1634285f-22, -4.307094f-19, 1.672701f-16, 1.2102875f-14, 5.515009f-14, 8.999296f-14, 2.8241575, -2.8241575]"

```

This results in:

```julia
ERROR: LoadError: ArgumentError: cannot parse " 3.5440563f-28" as Float32

```

What would be a good way to avoid this error?

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 4, 2024, 12:51pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/6 "2024-09-04T12:51:02Z")

</div>

> [@ufechner7](#):
>
> What would be a good way to avoid this error?

`str = replace(str, 'f'=>'e')`. You could also use [GitHub - JuliaData/Parsers.jl: fast parsing machinery for basic types in Julia](https://github.com/JuliaData/Parsers.jl), which should support this format directly.

See [`parse` can't parse our own Float32 output format · Issue #5690 · JuliaLang/julia · GitHub](https://github.com/JuliaLang/julia/issues/5690)

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [September 4, 2024, 12:51pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/7 "2024-09-04T12:51:39Z")

</div>

Why is `Float32[` in the CSV file? Wouldn’t it just be the actual float values as strings?

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 1:41pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/8 "2024-09-04T13:41:36Z")

</div>

> [@StefanKarpinski](#):
>
> Why is `Float32[` in the CSV file?

Because I exported it from a StructArray using CSV.jl

> <https://github.com/ufechner7/KiteUtils.jl/blob/93e5e5e20c5db1b9ba9e0e40c8d4426d7a26af95/src/KiteUtils.jl#L587C1-L593C4>

Correction:  
`syslog` is a struct array and some of the elements of the struct are Vectors.  
CSV.write is writing them in this way.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 1:55pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/9 "2024-09-04T13:55:26Z")

</div>

For future reference:

```julia
using Parsers

function parse_vector(str)
    m = match(r"\[(.*)\]", str)
    strs = split(m[1], ','; keepempty=false)
    Parsers.parse.(Float32, strs)
end

```

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [September 4, 2024, 5:00pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/10 "2024-09-04T17:00:18Z")

</div>

> [@ufechner7](#):
>
> Because I exported it from a StructArray using [CSV.jl](https://juliahub.com/ui/Packages/General/CSV)

If you are controlling the format, then I would use a different format from CSV for this kind of data. Something like JSON or JLD/HDF5 seems more suited.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 5:04pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/11 "2024-09-04T17:04:27Z")

</div>

> [@stevengj](#):
>
> If you are controlling the format, then I would use a different format

Well, by default I use the Apache arrow format. But it causes some issues when the format changes (I am adding or removing columns). In that case, CSV is a more generic choice that still works after I added columns. So if I have an old log file I export it as CSV and import it into the new format and then save it in the new format as .arrrow file.

I needed this function for this purpose.

And I am very happy to have a 2.5 dimensional log file format where columns can contain vectors. Easy to process and filter and export to Python, for example.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [September 4, 2024, 5:45pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/12 "2024-09-04T17:45:52Z")

</div>

> [@ufechner7](#):
>
> And I am very happy to have a 2.5 dimensional log file format where columns can contain vectors. Easy to process and filter and export to Python, for example.

Well, except for this problem. I think that @stevengj’s point is that in this case you’re saving more complex nested data structures, which makes CSV awkward.

If you’re going to have CSVs where some columns are arrays, but not get any more complex than that, it might be better to custom print those in a format that’s simpler to parse instead of relying on Julia’s value printing, e.g. represent a vector of values as a string containing delimiter-separated values.

If the structure will be arbitrarily nested and complex, then you’re going to keep hitting these issues and will definitely want a standard format like JSON or HDF5, both of which can be easily read in other languages, but it sounds like you’re not in that case.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 6:04pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/13 "2024-09-04T18:04:09Z")

</div>

Well, I used Google protocol buffer messages for our log files in the past to be fully n-dimensional, but in the end, if you want to plot or analyze the data you need 2-dimensional data with a fixed sampling rate.

My system is represented as particle system with a small number of particles, like 11, and I have columns for the X, Y and Z position vectors of these particles, and one column for the quaternium that represents the orientation of the kite.

No need for deeply nested structures.

And I like Julias way to print vectors: Easily readable and clear. You can round-trip it because the type is included. OK, the “f” instead of “e” for the exponent in scientific notation is somewhat confusing.

---

<div class="post-metadata">

**Author:** ![StefanKarpinski](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stefankarpinski/32/24_2.png) [@StefanKarpinski](https://discourse.julialang.org/u/StefanKarpinski)\
**Post date:** [September 4, 2024, 8:05pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/14 "2024-09-04T20:05:30Z")

</div>

> [@ufechner7](#):
>
> and one column for the quaternium that represents the orientation of the kite.

You could also explode that into multiple scalar columns.

---

<div class="post-metadata">

**Author:** ![ufechner7](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ufechner7/32/51363_2.png) [@ufechner7](https://discourse.julialang.org/u/ufechner7)\
**Post date:** [September 4, 2024, 8:41pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/15 "2024-09-04T20:41:02Z")

</div>

> [@StefanKarpinski](#):
>
> You could also explode that into multiple scalar columns.

Very bad idea. Well, for the orientation that would be an option, but for the position vectors not, because the number of elements varies. And I do not want to have a different number of columns depending on the simulation that I run. Finally, CSV is not the main format I use, the main format is Apache arrow, and in arrow format you can have vectors in a column in a very well defined way. And in the CSV format I want to have the same columns as in the arrow files, otherwise automatic conversion is not possible.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [September 4, 2024, 10:04pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/16 "2024-09-04T22:04:53Z")

</div>

> [@ufechner7](#):
>
> Very bad idea.

It is not so much a bad idea as a badly defined problem to begin with.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [September 4, 2024, 10:32pm UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/17 "2024-09-04T22:32:56Z")

</div>

Another option:

```julia
str = "Float32[0.0, 3.5440563f-28, -8.439492f-25, 7.1634285f-22, -4.307094f-19, 1.672701f-16, 1.2102875f-14, 5.515009f-14, 8.999296f-14, 2.8241575, -2.8241575]"

ix1, ix2 = findfirst('[', str) + 1, findfirst(']', str) - 1
parse.(Float32, split(replace(str[ix1:ix2], 'f'=>'e'), ','))

```

---

<div class="post-metadata">

**Author:** ![liuyxpp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/liuyxpp/32/9870_2.png) [@liuyxpp](https://discourse.julialang.org/u/liuyxpp)\
**Post date:** [September 5, 2024, 1:04am UTC](https://discourse.julialang.org/t/parse-string-representation-of-a-vector-into-floats/119025/18 "2024-09-05T01:04:30Z")

</div>

Assume your maximum number of elements of your position vectors is 3. Then no matter the actual number of elements of your position vectors in a specific simulation, I would save them as if they are 3D with other dim value being dummy values (usually 0?). It may increase the size of your CSV file for 1D or 2D simulation data, but at least they are consistent across all possible dimensions.
