# \[WIP\] Announcing Volcanito.jl: a backend-agnostic interface for tabular data

**URL:** <https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695>\
**Category:** Package Announcements\
**Created:** [August 28, 2020, 2:44pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695 "2020-08-28T14:44:03Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 28, 2020, 2:44pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/1 "2020-08-28T14:44:04Z")

</div>

I want to announce a public demo release of a package I’ve been working on to test the viability of building a backend agnostic interface for tabular data in Julia that would let the ecosystem move towards an architecture closer to that used by databases. In homage to the classic Volcano model in the database literature, I’ve called the package Volcanito and you can check it out at [GitHub - johnmyleswhite/Volcanito.jl: A backend agnostic for tabular data operations in Julia](https://github.com/johnmyleswhite/Volcanito.jl).

The package lets users write operations on data in terms of simple macros:

```julia
@select(df, a, b, d = a + b)

@where(df, a > b)

@aggregate_vector(
    @group_by(df, !c),
    m_a = mean(a),
    m_b = mean(b),
    n_a = length(a),
    n_b = length(b),
)

@order_by(df, a + b)

@limit(df, 10)

```

These macros are translated into logical nodes that can be applied to arbitrary data sources.

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 3:08pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/2 "2020-08-28T15:08:19Z")

</div>

Do you plan to add `@join`?

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 28, 2020, 3:13pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/3 "2020-08-28T15:13:43Z")

</div>

Yes, I’ll get there at some point, but I’ve only had a little bit of time to work on this over my last week of vacation, so it might take a long time to get there. The goal here was to mostly to get something whose architecture is mature enough out into the public to influence future thinking in the space.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 28, 2020, 3:15pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/4 "2020-08-28T15:15:23Z")

</div>

I have been thinking about something like this.

This is nice.

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 3:17pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/5 "2020-08-28T15:17:22Z")

</div>

How does this compare to [Query.jl](https://github.com/queryverse/Query.jl) in [Queryverse](https://github.com/queryverse)?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 28, 2020, 3:28pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/6 "2020-08-28T15:28:53Z")

</div>

Query.jl seems to be row-semantic based where as Volcanito.jl have greater potential as column-nar manipulation library.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 28, 2020, 3:30pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/7 "2020-08-28T15:30:17Z")

</div>

I see that it does lazy evaluation and only when `show` is hit does it show the results. This seems like it will hit bottlenecks if the datasets are large as neither caching nor recomputing on the fly would be good solutions

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 3:35pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/8 "2020-08-28T15:35:04Z")

</div>

Why would lazy evaluation hit bottlenecks for large datasets?

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 28, 2020, 3:55pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/9 "2020-08-28T15:55:55Z")

</div>

I don’t know. I’ve never made much effort to follow the Queryverse. I wrote this in part as an indication of how I would hope DataFramesMeta would evolve.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 28, 2020, 4:09pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/10 "2020-08-28T16:09:35Z")

</div>

> [@derekmahar](#):
>
> Why would lazy evaluation hit bottlenecks

Why wouldn’t it? Every time u print it runs thru the same operations. Each operation might take 10 mins. Unless cached. But cache three operation might be huge cos the data is huge.

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 4:09pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/11 "2020-08-28T16:09:47Z")

</div>

How would one compose these macros? For example, how would one compose `@select` and `@where` to [select or reject](https://discourse.julialang.org/t/the-filter-function-is-non-inntuitive/45674/6) certain rows which match the `@where` condition?

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [August 28, 2020, 4:12pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/12 "2020-08-28T16:12:45Z")

</div>

> [@derekmahar](#):
>
> example, how would one compose `@select` and `@where` select or reject certain rows which match the `@where` condition?

Based on my understanding it builds a DAG of sorts and compiles that DAG to DataFrames.jl code.

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 4:13pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/13 "2020-08-28T16:13:42Z")

</div>

> [@xiaodai](#):
>
> Why wouldn’t it? Every time u print it runs thru the same operations. Each operation might take 10 mins. Unless cached. But cache three operation might be huge cos the data is huge.

I follow what you mean. Why not, in addition to `show`, introduce an `@compute` or `@result` operations that calculate and cache an intermediate result to which the program could apply additional lazy operations?

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 28, 2020, 4:35pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/14 "2020-08-28T16:35:23Z")

</div>

They already should compose. The example depends on composition:

```julia

@aggregate_vector(
    @group_by(df, !c),
    m_a = mean(a),
    m_b = mean(b),
    n_a = length(a),
    n_b = length(b),
)

```

Your select and where example is equivalent; just replace `df` in one of the expressions with the result from another operation.

---

<div class="post-metadata">

**Author:** ![jlapeyre](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jlapeyre/32/4514_2.png) [@jlapeyre](https://discourse.julialang.org/u/jlapeyre)\
**Post date:** [August 28, 2020, 4:35pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/15 "2020-08-28T16:35:49Z")

</div>

> [@derekmahar](#):
>
> I follow what you mean. Why not, in addition to `show` , introduce an `@compute` or `@result` operations

If you read the docs, you see that this already exists: `materialize`.  
It looks like this is meant to allow a declarative query that can be optimized when executed. You are always free to execute the graph and then use the result in further operations.

---

<div class="post-metadata">

**Author:** ![johnmyleswhite](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/johnmyleswhite/32/31_2.png) [@johnmyleswhite](https://discourse.julialang.org/u/johnmyleswhite)\
**Post date:** [August 28, 2020, 4:35pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/16 "2020-08-28T16:35:56Z")

</div>

What’s the difference here from the `materialize` operation used in the second part of the README?

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 4:40pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/17 "2020-08-28T16:40:12Z")

</div>

> [@jlapeyre](#):
>
> If you read the docs, you see that this already exists: `materialize` .

Yes, you’re right. I overlooked the last sentence and last few lines of the second example. @xiaodai might have missed these, too. 😉

---

<div class="post-metadata">

**Author:** ![derekmahar](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/derekmahar/32/216439_2.png) [@derekmahar](https://discourse.julialang.org/u/derekmahar)\
**Post date:** [August 28, 2020, 4:40pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/18 "2020-08-28T16:40:51Z")

</div>

> [@johnmyleswhite](#):
>
> What’s the difference here from the `materialize` operation used in the second part of the README?

Nothing. They are the same idea. I just didn’t see it until @jlapeyre pointed it out to me. 🙂

---

<div class="post-metadata">

**Author:** ![juliohm](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/juliohm/32/215266_2.png) [@juliohm](https://discourse.julialang.org/u/juliohm)\
**Post date:** [August 28, 2020, 6:30pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/19 "2020-08-28T18:30:47Z")

</div>

Exactly what I am looking for @johnmyleswhite, thank you for the contribution. Could anyone please provide some comparison with the Query.jl package, what are the pros and cons of each approach?

Does any of them provide a row selector? I would like to slice Tables.jl tables vertically in a lazy fashion given indices for start and end rows, but couldn’t find a package to do this yet.

---

<div class="post-metadata">

**Author:** ![anon92994695](https://avatars.discourse-cdn.com/v4/letter/a/ce7236/32.png) [@anon92994695](https://discourse.julialang.org/u/anon92994695)\
**Post date:** [August 28, 2020, 7:51pm UTC](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695/20 "2020-08-28T19:51:20Z")

</div>

Fantastic idea! I was kind of hoping this would be Vulcanito.jl ([Vulcans | Star Trek](https://www.startrek.com/database_article/vulcans)) but you’re naming reasoning is much more sound.

What backends do you intend to support? DataKnots.jl?

[Next page](https://discourse.julialang.org/t/wip-announcing-volcanito-jl-a-backend-agnostic-interface-for-tabular-data/45695.md?page=2)
