# JuliaDB versus

**URL:** https://discourse.julialang.org/t/juliadb-versus/20977
**Category:** Data
**Tags:** question
**Created:** [February 19, 2019, 4:19pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977 "2019-02-19T16:19:27Z")
**Posts on this page:** 13
**Page:** 1

<div class="post-metadata">

### Author: ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)
#### Post date: [February 19, 2019, 4:19pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/1 "2019-02-19T16:19:27Z")

</div>

I’ve been using JuliaDB and absolutely loving it . It helped me restructure my data and I’m now able to process my data in about 40 LOC. It all makes sense.

In my application, I’m `join`ing and grouping sources together to distill a table that contains all the data needed for the final step. This final step is costly (each iteration – each row – takes a few seconds). What I’m missing though, is some sort of “piping”:  
I’m executing this costly step on each row with a `groupby` (so it groups the data and then applies the step). The result of which is the final product (the `groupby` returns a table when it’s done, not per row). But since I have hundreds of rows and each row is slow, stopping the process midway causes all the data to get lost (why stop the process one might ask, good question). What would be good is some `DataFrames.groupby` that iterates over the groups and has some side-effect (like saving or piping it to a sink). As I’m writing this, I figure I could just create a grouped table (that’s of course very fast), and then in a for-loop iterate over the rows, saving the results as I go. Yea, that’s basically the same.

OK, I’ll post this just in case someone has a better suggestion. Sorry for the somewhat vague post 😝

---

<div class="post-metadata">

### Author: ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)
#### Post date: [February 19, 2019, 5:07pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/2 "2019-02-19T17:07:43Z")

</div>

[Query.jl](https://github.com/queryverse/Query.jl) allows you to hook into the pipeline:

```julia
julia> df = DataFrame(a=[1,1,2,3,2], b=rand(5))
5×2 DataFrame
│ Row │ a │ b │
│ │ Int64 │ Float64 │
├─────┼───────┼──────────┤
│ 1 │ 1 │ 0.70583 │
│ 2 │ 1 │ 0.109909 │
│ 3 │ 2 │ 0.209058 │
│ 4 │ 3 │ 0.70213 │
│ 5 │ 2 │ 0.85946 │

julia> df |> 
          @groupby(_.a) |>
          @map(i-> begin                                                  
                  @info key(i)                                                                        
                  @info i                                                                             
                  return {a = key(i), b=sum(i.b)}                                                     
              end) |>
          DataFrame                                                                   
[ Info: 1                                                                                     
[Info: NamedTuple{(:a, :b),Tuple{Int64,Float64}}[(a = 1, b = 0.70583), (a = 1, b = 0.109909)]
[ Info: 2                                                                                     
[Info: NamedTuple{(:a, :b),Tuple{Int64,Float64}}[(a = 2, b = 0.209058), (a = 2, b = 0.85946)]
[ Info: 3                                                                                     
[Info: NamedTuple{(:a, :b),Tuple{Int64,Float64}}[(a = 3, b = 0.70213)]                       
3×2 DataFrame                                                                                 
│ Row │ a │ b │                                                                    
│ │ Int64 │ Float64 │                                                                    
├─────┼───────┼──────────┤                                                                    
│ 1 │ 1 │ 0.815739 │                                                                    
│ 2 │ 2 │ 1.06852 │                                                                    
│ 3 │ 3 │ 0.70213 │                                                                    

```

The `@groupby` clause iterates `Grouping`s. A `Grouping` is an `AbstractArray` that holds the rows that belong to that group, plus you can call `key(g)` on the group to retrieve the key of that group.

In theory `@map` should normally be side-effect free, but I guess there is no real harm in adding some diagnostic output like here.

---

<div class="post-metadata">

### Author: ![piever](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/piever/32/1815_2.png) [@piever](https://discourse.julialang.org/u/piever)
#### Post date: [February 19, 2019, 5:38pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/3 "2019-02-19T17:38:35Z")

</div>

You can probably also do `groupby(identity, t, by; select = ...)` to get a table where one column has the grouped tables and then work on it. This should be reasonably fast as you are just taking views of the original.

---

<div class="post-metadata">

### Author: ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)
#### Post date: [February 19, 2019, 6:42pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/4 "2019-02-19T18:42:01Z")

</div>

> [@piever](#):
>
> do `groupby(identity, t, by; select = ...)` to get a table where one column has the grouped tables and then work on it.

Yea, that’s what I figured as well. I’ll try it out now 🙂

---

<div class="post-metadata">

### Author: ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)
#### Post date: [February 19, 2019, 8:08pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/5 "2019-02-19T20:08:25Z")

</div>

Maybe try out [LightQuery](https://bramtayl.github.io/LightQuery.jl/latest/) (preferably the dev version)

---

<div class="post-metadata">

### Author: ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)
#### Post date: [February 19, 2019, 8:11pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/6 "2019-02-19T20:11:26Z")

</div>

> [@bramtayl](#):
>
> Maybe try out [LightQuery](https://bramtayl.github.io/LightQuery.jl/latest/)

Yea, I was think that too!

---

<div class="post-metadata">

### Author: ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)
#### Post date: [February 19, 2019, 8:13pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/7 "2019-02-19T20:13:50Z")

</div>

If you try it out let me know how it goes esp. in terms of performance.

---

<div class="post-metadata">

### Author: ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)
#### Post date: [February 19, 2019, 8:15pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/8 "2019-02-19T20:15:18Z")

</div>

> [@bramtayl](#):
>
> let me know how it goes esp. in terms of performance

👍

---

<div class="post-metadata">

### Author: ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)
#### Post date: [February 22, 2019, 9:37pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/9 "2019-02-22T21:37:30Z")

</div>

Updates?

---

<div class="post-metadata">

### Author: ![yakir12](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/yakir12/32/297_2.png) [@yakir12](https://discourse.julialang.org/u/yakir12)
#### Post date: [February 27, 2019, 2:42pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/10 "2019-02-27T14:42:28Z")

</div>

Yes! So I ended up creating a new table with `groupby`, as I mentioned in the beginning, and haven’t YET tried your LightQuery. As an aside, it’s a bit overwhelming to test all the different query-like APIs out there, so things take time, a lot of time. One thing is for sure, the more packages I try the better my data becomes… Funny.

---

<div class="post-metadata">

### Author: ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)
#### Post date: [February 27, 2019, 6:50pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/11 "2019-02-27T18:50:07Z")

</div>

Np. Just trying to pawn the hard work of bench-marking onto someone else.

---

<div class="post-metadata">

### Author: ![versipellis](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/versipellis/32/7878_2.png) [@versipellis](https://discourse.julialang.org/u/versipellis)
#### Post date: [June 17, 2019, 7:30pm UTC](https://discourse.julialang.org/t/juliadb-versus/20977/12 "2019-06-17T19:30:52Z")

</div>

@bramtayl I’m curious, do you know how well LightQuery works with out-of-core JuliaDB datasets? Per [the docs](https://juliacomputing.github.io/JuliaDB.jl/latest/out_of_core/), there’s a much more limited subset of processes that can be run once it’s out-of-core, and I’m not entirely sure where LightQuery sits in all of the Query frameworks out there.

---

<div class="post-metadata">

### Author: ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)
#### Post date: [June 18, 2019, 12:35am UTC](https://discourse.julialang.org/t/juliadb-versus/20977/13 "2019-06-18T00:35:48Z")

</div>

@versipellis I haven’t tried. LightQuery is in a bit of a state of flux right now, as I’m simultaneously adding SQL support to Query and LightQuery. But I suspect the answer will be yes, provided that

1. data is pre-sorted
2. you can index out of order
