# \[WIP\] Announcing Volcanito.jl: a backend-agnostic interface for tabular data

**URL:** <https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695>\
**Category:** Package Announcements\
**Created:** [August 28, 2020, 2:44pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695 "2020-08-28T14:44:03Z")\
**Posts on this page:** 5\
**Page:** 3

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 30, 2020, 1:21am UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/41 "2020-08-30T01:21:31Z")

</div>

```julia
> df <- data.frame(x = c(1, 5, 6))
> mutate(df, deviation = x - mean(x))
  x deviation
1 1 -3
2 5 1
3 6 2

```

Agreed that it’s useful to have some sugar for doing this beyond the explicit aggregation and joining. That’s obviously what inspired SQL to add WINDOW as a syntactic feature 🙂

What makes me so averse to `x - mean(x)` is that it makes it impossible to do something like the following with any automatic parallelism:

```julia
df <- data.frame(x = c(1, 1, 2, 3))
mutate(df, foo = x + length(unique(x)))

```

The problem is that a composed function like `length(unique(x))` isn’t tractable to automatic parallelization. If the system has to always assume non-parallelized functions might be called it can either decide that (a) users need to explicitly ensure all functions they use explicitly describe their parallelization strategy (which is how distributed DB’s like Presto work when users add new aggregation functions) or it can choose to (b) never provide automatic parallelization. I think the latter approach is a pretty bad long-term bet for the foreseeable future of computing hardware.

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 30, 2020, 1:31pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/42 "2020-08-30T13:31:04Z")

</div>

As a side question, do you think it is possible to define these operations taking into account metadata? Say I have additional information like spatial coordinates, timestamps, etc for each row of the table and that this information lives outside the table object. How we can apply a `@groupby` for example and retain these metadata in the results? It would be nice if these efforts could handle these use cases.

The simplest approach would be to operate on the indices of the rows and introduce intermediate functions like `@groupby_inds` that developers could leverage to extract the indices of both the table and the metadata.

---

<div class="post-metadata">

**Author:** ![dpsanders](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dpsanders/32/3573_2.png) [@dpsanders](https://discourse.julialang.org/u/dpsanders)\
**Post date:** [August 30, 2020, 2:52pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/43 "2020-08-30T14:52:55Z")

</div>

> [@dlakelan](#):
>
> Soon someone will write a sql\_str macro

There’s [GitHub - wookay/Octo.jl: Octo.jl 🐙 is an SQL Query DSL in Julia](https://github.com/wookay/Octo.jl)

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 30, 2020, 3:00pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/44 "2020-08-30T15:00:16Z")

</div>

> As a side question, do you think it is possible to define these operations taking into account metadata? Say I have additional information like spatial coordinates, timestamps, etc for each row of the table and that this information lives outside the table object. How we can apply a `@groupby` for example and retain these metadata in the results? It would be nice if these efforts could handle these use cases.

I may not be understanding, but this seems either (a) like it should fall out naturally of normal SQL-style operations or (b) is a bit of a niche use case. Hard to imagine I’ll have enough spare time to ever get to (b).

---

<div class="post-metadata">

**Author:** ![kristoffer.carlsson](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/kristoffer.carlsson/32/22_2.png) [@kristoffer.carlsson](https://discourse.julialang.org/u/kristoffer.carlsson)\
**Post date:** [August 30, 2020, 3:13pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/45 "2020-08-30T15:13:01Z")

</div>

> [@dpsanders](#):
>
> > [@dlakelan](#):
> >
> > Soon someone will write a sql\_str macro
> 
> There’s [https://github.com/wookay/Octo.jl](https://github.com/wookay/Octo.jl)

The amazing thing is that this is not a `_str` macro!

[Previous page](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695.md?page=2)
