# Parquet: writing data as row groups

**URL:** <https://discourse.julialang.org/t/parquet-writing-data-as-row-groups/117247>\
**Category:** Data\
**Tags:** question\
**Created:** [July 19, 2024, 5:45pm UTC](https://discourse.julialang.org/t/parquet-writing-data-as-row-groups/117247 "2024-07-19T17:45:36Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kmiernik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kmiernik/32/202783_2.png) [@kmiernik](https://discourse.julialang.org/u/kmiernik)\
**Post date:** [July 19, 2024, 5:45pm UTC](https://discourse.julialang.org/t/parquet-writing-data-as-row-groups/117247/1 "2024-07-19T17:45:36Z")

</div>

I would appreciate help with my problem. I have data that come in series. After processing raw input I have about 100 columns and about 10^8 rows per series (and about 100 series in total). So far I have been using HDF5 format to store it, but I would like to test if parquet performs better, especially if I select later rows by filtering some specific column values.

I have imagined a parquet file with row groups, one for each series (the columns and the data format will be always the same, but each series might have been taken with different conditions, so the data interpretation may differ). I have tried the Parquet2.jl module, but I’m kind of stuck with how to achieve row groups in a file (or directory). Here’s a piece of code that generates some pseudo-data

```julia
function generate_data(N)
    Mmax = 100
    hits = zeros(UInt16, Mmax * 2 + 1, N)

    for i in 1:N
        M = rand(1:10)
        hits[1, i] = UInt16(M)
        dets = sample(1:Mmax, M, replace=false)
        for d in dets
            hits[2*(d-1)+2, i] = round(UInt16, rand() .* 30000.0, RoundDown)
            t = randn() .* 1000.0 + 5000.0
            if t < 0
                t = 0.0
            end
            hits[2*(d-1)+3, i] = round(UInt16, t)
        end
    end

    data = (M=hits[1, :], )
    names = ["E", "t"]
    for j in 1:Mmax
        for k in 1:2
            data = merge(data, (Symbol("$(names[k])_$j") => hits[2*(j-1)+k+1, :], ))
        end
    end
    
    data
end

```

And this function I’ve tried to append a series to a file

```julia
function write_parquet(filename; k=10, N=1_000_000)
    open(filename, write=true) do io
        fw = Parquet2.FileWriter(io)
        for i in 1:k
            data = generate_data(N)
            Parquet2.writetable!(fw, data) 
        end
        Parquet2.finalize!(fw)
    end
end

```

The resulting file for k=1 has about 82M, and for k=10, 820M, so it seems that the data are written. But upon opening

```julia
ds = Parquet2.Dataset(filename)

```

there is only one group, with 10^6 rows.

In principle it should be possible to append data to unfinished files (e.g. [fastparquet write function](https://fastparquet.readthedocs.io/en/latest/api.html#fastparquet.write) has “append” keyword that allows to add new rowgroup), but can it be achieved with Parquet2.jl (or Parquet.jl)?

---

<div class="post-metadata">

**Author:** ![kmiernik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kmiernik/32/202783_2.png) [@kmiernik](https://discourse.julialang.org/u/kmiernik)\
**Post date:** [July 22, 2024, 11:38am UTC](https://discourse.julialang.org/t/parquet-writing-data-as-row-groups/117247/2 "2024-07-22T11:38:52Z")

</div>

My mistake, the above function works properly. Upon opening the resulting file, there are _k_ groups

```julia
ds = Parquet2.Dataset("test.prq")
Parquet2.nrowgroups(ds)

```

and _ds_ can be iterated over.
