# Exact numerical logging

**URL:** <https://discourse.julialang.org/t/exact-numerical-logging/58766>\
**Category:** General Usage\
**Tags:** question\
**Created:** [April 7, 2021, 3:25pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766 "2021-04-07T15:25:34Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [April 7, 2021, 3:25pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/1 "2021-04-07T15:25:34Z")

</div>

I have a long-running calculation, for which want to emit a log that allows recovery of _exact_ numerical values, so that I can restart or debug the calculation. The typical logged item can be thought of as a (possibly nested) `NamedTuple` (or `Dict`), with vectors and scalars, eg

```julia
(a = rand(SVector{3,Float64}, 4), μ = 9.7, κ = rand(9))

```

It is not a problem if logging does not preserve types (eg the `SVector` above). The output would ideally be human-readable, reasonably future-proof, and support UTF8.

Suggestions on how to implement this would be helpful, even if very general. I am thinking of writing to a JSON file, but if someone already implemented something like this, I would love to hear about it.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 7, 2021, 3:59pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/2 "2021-04-07T15:59:21Z")

</div>

> [@Tamas\_Papp](#):
>
> I have a long-running calculation, for which want to emit a log that allows recovery of _exact_ numerical values, so that I can restart or debug the calculation.

What’s wrong with the `print` output (as used in JSON or whatever)? For `Float64` values, `print` outputs a decimal representation that, when parsed as `Float64`, yields exactly the same value.

Using your example:

```julia
julia> x = (a = rand(SVector{3,Float64}, 4), μ = 9.7, κ = rand(9));

julia> eval(Meta.parse(repr(x))) == x
true

```

---

<div class="post-metadata">

**Author:** ![stephenll](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stephenll/32/28751_2.png) [@stephenll](https://discourse.julialang.org/u/stephenll)\
**Post date:** [April 7, 2021, 4:48pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/3 "2021-04-07T16:48:26Z")

</div>

The professional developers on my team using c#, convert the floats (or Ints or bools), to an unsigned integers, then use base64 encoding. I’m told this will guarantee what is written out will be read back in, assuming using a windows or Mac computer. We don’t use Linux. We store everything as vector, and whatever format (json, xml, TOML) we are using also store dimensions and the datatype of the original data. Below is code I use to convert a vector to that hashed format and reading it back into a vector.

We’ve had a lot of success using TOML to store thousands of regression and unit tests across our systems using this. It made it very easy for me to write Julia versions of our c# and guarantee that I get the same results as our c#.

In the TOML, we store the binary representation of the data, but also the values written out. We do that so that from inspecting the file you can see the values, but if you want the exact information the program would use the binary rep. TOML makes it very easy to convert to Dict and back. In fact, we round trip TOML to Dict to Julia Structs.

If a vector, the size is a single number. If the data is a scalar, size is not included in the input.

We prefer TOML since it understands nan, infinity, easily maps to Dicts, human readable, and a lot of libraries are available.

The TOML file for one of the inputs looks like this:

```julia
[input.processrisk]
size = [2,3]
dtype = "Float64"
data = [0.09618765866371004, 0.44433488489147255, 0.9005311838970897, 0.6818201970996056, 0.8526313332776663, 0.41132839089506557]
bdata = "UEw9IMGfuD/QK8WV+2/cP5q3+8Um0ew/THEJl3jR5T9WHn+BwUjrPwzhs1A0U9o/"

```

Here is the Julia code with an example for working with a vector to create what we call the binary representation:

```julia
x = rand(10)

str = writebinarydata(x)

x2 = readbinarydata(str, Float64)

x .== x2

function writebinarydata(x::AbstractVector)

    if eltype(x) == Float64
		type = UInt64
    elseif eltype(x) == Bool
    	type = UInt8
    elseif eltype(x) == Int64
    	type = UInt64
    else 
    	@warn "Unsupported type"
    end

    reinterpret.(type,x) |> base64encode
end

readbinarydata(x::String, dtype::Type) = dtype.(reinterpret(dtype, x |> base64decode))

```

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [April 7, 2021, 5:47pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/4 "2021-04-07T17:47:06Z")

</div>

Is this guaranteed? Because I remember having this problem with C++ in which the original numbers and the ones printed and parsed back differed, and I had to start using `"%a"` format of `printf` to print the numbers in the hexadecimal representation (with this one C/C++ had no loss of precision).

---

<div class="post-metadata">

**Author:** ![adolgert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adolgert/32/20286_2.png) [@adolgert](https://discourse.julialang.org/u/adolgert)\
**Post date:** [April 7, 2021, 5:58pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/5 "2021-04-07T17:58:27Z")

</div>

Hi Tamas. I understand why it’s nice to have a human-readable file. If you can relax that requirement, the HDF file format is amazing for both preservation of exact values and for archival reliability. It underlies the JLD data format, too. HDF is exact about the binary representation of data values. It has built-in tools that will dump values to ASCII on-demand, so it’s sort of as good as ascii? And it’s supported by a non-profit since 1988, or so. It’s kind of great, when its complexity is tamed by an API layer like JLD.

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 7, 2021, 6:19pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/6 "2021-04-07T18:19:06Z")

</div>

Not sure if this counts as human readable, but Julia is capable of printing and parsing Floats in Base16 which removes the room for interpretation in how they get parsed.

---

<div class="post-metadata">

**Author:** ![rafael.guerra](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/rafael.guerra/32/216610_2.png) [@rafael.guerra](https://discourse.julialang.org/u/rafael.guerra)\
**Post date:** [April 7, 2021, 6:32pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/7 "2021-04-07T18:32:40Z")

</div>

NB: @Oscar_Smith, to understand your post better, [this](https://stackoverflow.com/questions/21882862/purpose-of-hexadecimal-floating-point-value) helped.

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 7, 2021, 6:39pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/8 "2021-04-07T18:39:53Z")

</div>

> [@Henrique\_Becker](#):
>
> Is this guaranteed?

Yes, it’s guaranteed by modern float-to-string algorithms like [Ryu](https://dl.acm.org/doi/10.1145/3192366.3192369) (which is employed by Julia) and Grisu (which Julia used to employ).

This “information preservation” property was codified by [Steele and White (1990)](https://dl.acm.org/doi/10.1145/93548.93559) and I think it’s become widely accepted.

(And the good news is that it’s implemented in pure Julia, so you don’t have to worry about what your operating system’s libc does.)

---

<div class="post-metadata">

**Author:** ![Henrique\_Becker](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/henrique_becker/32/15443_2.png) [@Henrique\_Becker](https://discourse.julialang.org/u/Henrique_Becker)\
**Post date:** [April 7, 2021, 6:50pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/9 "2021-04-07T18:50:34Z")

</div>

Very good to know. A shame this is not adopted by more languages. I had this problem about 4 years ago, so it seems like C/C++ is not set on adopting it (as there were C/C++ standard updates between 1990 and 2016 and the problem still existed at 2016).

---

<div class="post-metadata">

**Author:** ![Oscar\_Smith](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oscar_smith/32/25343_2.png) [@Oscar\_Smith](https://discourse.julialang.org/u/Oscar_Smith)\
**Post date:** [April 7, 2021, 6:51pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/10 "2021-04-07T18:51:30Z")

</div>

The problem here is that C/C++ farm a lot out the the OS, and some OSes are pretty bad with some of the finicky stuff like this.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [April 8, 2021, 6:59am UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/11 "2021-04-08T06:59:58Z")

</div>

> [@stevengj](#):
>
> What’s wrong with the `print` output (as used in JSON or whatever)?

Nothing, that’s fine. I just need to embed it in JSON etc so that it is easier to process the file with a script if necessary.

> [@stephenll](#):
>
> We’ve had a lot of success using TOML to store thousands of regression and unit tests across our systems using this.

Does TOML (and its current Julia implementation) scale to large files?

> [@adolgert](#):
>
> the HDF file format is amazing for both preservation of exact values and for archival reliability.

Yes, HDF is fine and I am using it elsewhere, but this data is really unstructured. The intention is to replace `@info` etc logging with something more structured and searchable.

---

<div class="post-metadata">

**Author:** ![stephenll](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stephenll/32/28751_2.png) [@stephenll](https://discourse.julialang.org/u/stephenll)\
**Post date:** [April 8, 2021, 11:32am UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/12 "2021-04-08T11:32:52Z")

</div>

The largest file we’ve created contains about 1 million floats in total. That takes about .4 seconds to write and about .3 seconds to read on my 2019 MacBook Pro.

---

<div class="post-metadata">

**Author:** ![Skoffer](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/skoffer/32/378_2.png) [@Skoffer](https://discourse.julialang.org/u/Skoffer)\
**Post date:** [April 8, 2021, 12:07pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/13 "2021-04-08T12:07:22Z")

</div>

Not sure, how close it to your question, but I think there is no need in removing `@info`, you just need to change your sink. For example [LoggingFacilities.jl](https://github.com/tk3369/LoggingFacilities.jl#jsontransformerlogger) provides conversion of usual logging output to `JSON` format. With the help of [LoggingExtras.jl](https://github.com/oxinabox/LoggingExtras.jl) you can setup your logging to write to HDF/BSON/whatever without changing a line of your code (yet initial logger setup take some time of course).

---

<div class="post-metadata">

**Author:** ![adolgert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adolgert/32/20286_2.png) [@adolgert](https://discourse.julialang.org/u/adolgert)\
**Post date:** [April 8, 2021, 4:18pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/14 "2021-04-08T16:18:49Z")

</div>

Tamas, what a nice idea! It’s a logstash-style of program serialization. I remain a cheerleader for HDF, but they do refer to strings as “Other Non-numeric Datatypes” in their user manual, so maybe there’s a little neglect in that area.

---

<div class="post-metadata">

**Author:** ![oxinabox](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/oxinabox/32/206603_2.png) [@oxinabox](https://discourse.julialang.org/u/oxinabox)\
**Post date:** [April 8, 2021, 4:41pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/15 "2021-04-08T16:41:48Z")

</div>

One optition is using TensorBoardLogger.jl.  
It logs to a ProtoBuf file, that can be displayed by the TensorBoard program,  
but also can be read by a number of libraries (including TensorBoardLogger.jl itself)

> **[GitHub - JuliaLogging/TensorBoardLogger.jl: Easy peasy logging to TensorBoard...](https://github.com/JuliaLogging/TensorBoardLogger.jl)**
>
> Easy peasy logging to TensorBoard with Julia. Contribute to JuliaLogging/TensorBoardLogger.jl development by creating an account on GitHub.

In general though i think we should make a bunch of what LoggingExtras.jl calls Logging Sinks.  
A JSON one would be good (lots of tools out there speak JSON for log files, e.g. AWS’s cloudwatch).  
I am not sure if one exists (other than the [LoggingFacilities.JSONTransformerLogger](https://github.com/tk3369/LoggingFacilities.jl#jsontransformerlogger) which is an unusual take on it, in that it isn’t a sink, it is a transformer, i see some advantages to that though).

Perhaps more generically would be good to make a Table logging sink.  
Which is setup to be able to generically write to any Tables.jl sink (e.g. a CSV, LibPQ, Arrow)

---

<div class="post-metadata">

**Author:** ![stevengj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/stevengj/32/71_2.png) [@stevengj](https://discourse.julialang.org/u/stevengj)\
**Post date:** [April 8, 2021, 4:43pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/16 "2021-04-08T16:43:20Z")

</div>

There is also the option of (Google’s) protocol buffers, which have a Julia implementation ([ProtoBuf.jl](https://github.com/JuliaIO/ProtoBuf.jl)). I don’t have any experience with them myself, but they are a binary-only JSON competitor and might be worth looking into for scaling to large sizes.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [April 13, 2021, 2:00pm UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/17 "2021-04-13T14:00:27Z")

</div>

After experimenting with many options, I settled on HDF5 after all, because it is the most robust and hassle-free. In particular, JSON & similar flattens all arrays to vectors, and ProtoBuf and TensorBoard are a bit heavyweight for my purposes.

I wrapped up everything in a mini-package

> **[GitHub - tpapp/HDF5Logging.jl: Log to HDF5 files from Julia.](https://github.com/tpapp/HDF5Logging.jl)**
>
> Log to HDF5 files from Julia.

which is currently being registered. Particular attention is paid to thread safety, not keeping files open when not needed, and reading back logs. Comments, PRs, etc are of course welcome.

# Example

```julia-repl
julia> using HDF5Logging, Logging

julia> logger = hdf5_logger(tempname())
Logging into HDF5 file /tmp/jl_IbbUvj, 0 messages in “log”

julia> # write log

julia> with_logger(logger) do
       @info "very informative" a = 1
       end

julia> # read log

julia> logger[1]
(level = Info, message = "very informative", _module = "Main", group = "REPL[46]", id = "Main_7a40b9cc", file = "REPL[46]", line = 2, data = ["a" => 1])

```

---

<div class="post-metadata">

**Author:** ![adolgert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/adolgert/32/20286_2.png) [@adolgert](https://discourse.julialang.org/u/adolgert)\
**Post date:** [April 15, 2021, 12:00am UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/18 "2021-04-15T00:00:20Z")

</div>

Tamas, this code uses metadata keys to store information attached to groups in the HDF5 file. This kind of key isn’t kept in block storage, so I’m curious whether it behaves well for even small numbers of logging messages. It’s like that Reeses commercial, “You put your data in my metadata!” How large did you test?

HDF5 has two main ways to store logging-type information: either a Packet Table or a Dataset with a Compound Datatype. You could also use an Opaque datatype (h5ex\_t\_opaque) to kind of stuff whatever you want in there.

The files also aren’t set up for an append operation. HDF5 stores data in blocks that are indexed by B+ trees and kept in an LRU cache. Packets and logging tend to write in chunks of messages, as a result, in order to reduce churn of blocks and to avoid corruption. Some of the other file formats mentioned are stars at appending and can recover from failed or partial writes.

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [April 15, 2021, 8:39am UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/19 "2021-04-15T08:39:16Z")

</div>

> [@adolgert](#):
>
> You put your data in my metadata!” How large did you test?

Yes, I am aware of the fact that this is a compromise. It works fine for me though as I don’t have a lot of log messages, just a few with large objects.

I would love to use JSON and TOML, but my current obstacle is serializing arrays without losing the dimension information. This basically means that I would have to invent my own format within either to keep track of this. Cf

> <https://github.com/JuliaData/StructTypes.jl/issues/14>
>
> Currently \`ArrayType()\` iterates element-wise over an array type using the defau…lt \`iterate(x::T)\` method when serializing. This has the effect of flattening higher dimensional arrays into vectors. Matrices, for instance, become vectors with no information about their original dimensions.
> 
> The JSON standard maintains array order, so it seems safe for me to serialize matrices as vectors of vectors. However, my only way to affect serialization is by overloading the \`iterate(x::T)\` function and doing so for the \`Matrix\` type seems likely to recompile large parts of Julia base (as well as break my own code.)
> 
> Then on deserialization, I would just \`construct(::Type{Matrix}, x::Vector) = hcat(x...)\`
> 
> It makes sense on some level that this would be natively unsupported because it makes the serialized representation of matrices and vectors of vectors indistinguishable (and thus \`deserialize(serialize(matrix)) != matrix\` without an additional construct definition), but that code doesn't work right now either and it seems like there should be some way to for users to enable support for it because matrices are so common. Perhaps there should be something analogous to \`construct()\` for serialization?

HDF5, even if slightly abused, seems like a reasonable workaround that allows me to proceed with the actual computation without getting lost in a rabbit hole.

In the long run, it would be great to have a protocol that just serializes everything to

1. integers and floats,
2. vectors of the above,
3. dictionaries of dictionaries and/or the above.

This could then be used with JSON, TOML, BSON, etc.

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [April 15, 2021, 9:14am UTC](https://discourse.julialang.org/t/exact-numerical-logging/58766/20 "2021-04-15T09:14:28Z")

</div>

> [@Tamas\_Papp](#):
>
> I have a long-running calculation, for which want to emit a log that allows recovery of _exact_ numerical values, so that I can restart or debug the calculation.

If this is the use case, I would just use the `Serialization` standard library.

[Next page](https://discourse.julialang.org/t/exact-numerical-logging/58766.md?page=2)
