# Avro.jl : Problems appending and reading one record at a time

**URL:** <https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [July 31, 2021, 3:23pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604 "2021-07-31T15:23:36Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jonm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jonm/32/27392_2.png) [@Jonm](https://discourse.julialang.org/u/Jonm)\
**Post date:** [July 31, 2021, 3:23pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/1 "2021-07-31T15:23:36Z")

</div>

I want to serial write a sequence of Dict{String,Int64} objects to disc in Avro ad they are generated. Hence I need to beable to close and then append to the file. Hence Avro (if I understand) seems like a good choice.  
I am using Avro.jl

I can’t seem to get the regular Avro.read to read more than the first record I write with Avro.write.

In the end I got this working by using Avro.readvalue which I found by reading the source code.

```julia
buff=Base.read(io)
obj, pos = Avro.readvalue(Avro.Binary(),schema,JuliaType,buff, 
            pos, length(buff),false)

```

the problem is that my files are quite large. around 5-8G. I really don’t want to read everything into memory into buff.

I can’t get the Avro.writetable working either without passing through DataFrames and that seems crazy. Also, I still have problems appending records.

Can I get Avro.read to read just the first Avro record and then read the next one the next time it is called ?

This works great in python. I am relatively new to julia.

Thanks in advance.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [August 1, 2021, 9:54pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/2 "2021-08-01T21:54:35Z")

</div>

I think @quinnj may be the relevant expert here.

> [@Jonm](#):
>
> I need to beable to close and then append to the file. Hence Avro (if I understand) seems like a good choice.

I wonder if Arrow Stream might also be a good choice.

---

<div class="post-metadata">

**Author:** ![Jonm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jonm/32/27392_2.png) [@Jonm](https://discourse.julialang.org/u/Jonm)\
**Post date:** [August 2, 2021, 9:41pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/3 "2021-08-02T21:41:11Z")

</div>

Thanks. I think I am switching to writing JSON’s line by line and compressing the whole file using a gzip stream. But now I am having issues with running tuples through JSON3

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [August 2, 2021, 11:46pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/4 "2021-08-02T23:46:05Z")

</div>

It sounds like avro can be a good data format for your use-case here. Currently, Avro.jl only supports writing multiple blocks like you suggested via the `Avro.writetable`/`Avro.readtable` interfaces. I’m not sure how you were writing to the file yourself, but I wouldn’t expect `Avro.readvalue` to then work at reading an arbitrary # of records. The avro format has the explicit “block” construct to allow arbitrary appending of record blocks in files, but it’s not really supported writing yourself.

One thing we could probably do is provide a more transactional way of writing, where you have an explicit “start writing”, “append N blocks of records”, “close writing” functions. If that would be useful, feel free to open an issue and I can look into it. We have something similar like this for `Arrow.write` already (allowing you to append a recordbatch to an existing file).

Alternatively, `Avro.writetable` _does_ support the `Tables.partitions` interface on inputs, so you could restructure your writing so the dicts are generated via `Tables.partitioner`, something like:

```julia
Avro.writetable(
    "data.avro",
    Tables.partitioner(1:N) do i
        # generate dict here
        # but need to return a valid "table", simplest is Vector{NamedTuple}
    )
)

```

This means that each “partition” (generated dict) will be written out to the file, one at a time.

You can then read back in the data by doing `tbl = Avro.readtable("data.avro")`, which uses a lazy reading, so even larger-than-RAM files should be fine.

---

<div class="post-metadata">

**Author:** ![quinnj](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/quinnj/32/11_2.png) [@quinnj](https://discourse.julialang.org/u/quinnj)\
**Post date:** [August 2, 2021, 11:46pm UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/5 "2021-08-02T23:46:27Z")

</div>

Feel free to open a JSON3.jl issue or post here with details; happy to help brainstorm the best approach here too.

---

<div class="post-metadata">

**Author:** ![jzr](https://avatars.discourse-cdn.com/v4/letter/j/eb9ed0/32.png) [@jzr](https://discourse.julialang.org/u/jzr)\
**Post date:** [August 3, 2021, 12:34am UTC](https://discourse.julialang.org/t/avro-jl-problems-appending-and-reading-one-record-at-a-time/65604/6 "2021-08-03T00:34:00Z")

</div>

For @quinnj and future readers following along, that JSON3 discussion is in

> [@JSON3 and Tuples](https://discourse.julialang.org/t/json3-and-tuples/65699/3):
>
> Thanks. That explains it. I actually have this dictionary as part of a struct (in code someone else wrote so can’t modify that). Is there some way to teach JSON3 to convert this to a [key, value] array and then when it reads it back in convert it to a dictionary. I could do this by hand, but it would be nice if it did this automagically. Thanks for the help again. I am new to using JSON
