# ANN: LightQuery tutorial

**URL:** <https://discourse.julialang.org/t/ann-lightquery-tutorial/20262>\
**Category:** Package Announcements\
**Tags:** tutorials\
**Created:** [January 30, 2019, 5:30am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262 "2019-01-30T05:30:52Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 5:30am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/1 "2019-01-30T05:30:52Z")

</div>

I’ve put out a 1.0 version of LightQuery, and with it I’ve done a tutorial. This tutorial matches the one for dplyr as closely as possible. See [https://bramtayl.github.io/LightQuery.jl/latest/#Tutorial-1](https://bramtayl.github.io/LightQuery.jl/latest/#Tutorial-1) for the tutorial.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 30, 2019, 5:41am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/2 "2019-01-30T05:41:23Z")

</div>

Suggests you rename your package to “MyPreciousQuery” given your tag line is “One query to rule them all” 😀

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 30, 2019, 5:46am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/3 "2019-01-30T05:46:13Z")

</div>

What’s this package for? Looks like it’s a row-oriented querying library? Which also works on column-oriented data. But the default data is in row-oriented format.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 5:51am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/4 "2019-01-30T05:51:19Z")

</div>

I’d definitely consider MyPreciousQuery. Some of the functions work with iterators (e.g. rows) and some work with NamedTuples (e.g. columns). You have to explicitly move your data back and forth between row-wise and columnwise using the `rows` and `columns` functions provided, depending on what you want to do.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 30, 2019, 5:58am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/5 "2019-01-30T05:58:54Z")

</div>

I thought the trend is towards columnar database because it’s generally faster due to the fact that in most cases you would want to apply functions to every element in the same column.

I was surprised that the default seems to be row-oriented? I can see how that will have bad performance for 100 million rows.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 6:03am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/6 "2019-01-30T06:03:25Z")

</div>

You can keep your data stored as columns if you want; I’d definitely recommend it. In fact, the flights data that I’m working with in the tutorial is stored as columns.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 6:04am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/7 "2019-01-30T06:04:31Z")

</div>

`rows` is a lazy iterator, and doesn’t affect how the underlying data is stored, if that clears things up.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 6:12am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/8 "2019-01-30T06:12:30Z")

</div>

A lot of the under-the-hood magic that allows column-wise collection is driven by unzip, which I’m particularly proud of.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 30, 2019, 7:46pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/9 "2019-01-30T19:46:43Z")

</div>

Genuinely curious about people’s reaction to this, esp. in comparison to other querying packages.

---

<div class="post-metadata">

**Author:** ![Yifan\_Liu](https://avatars.discourse-cdn.com/v4/letter/y/4da419/32.png) [@Yifan\_Liu](https://discourse.julialang.org/u/Yifan_Liu)\
**Post date:** [January 30, 2019, 9:32pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/10 "2019-01-30T21:32:10Z")

</div>

Is this package any different from Query.jl? Just curious.

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 30, 2019, 9:45pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/11 "2019-01-30T21:45:07Z")

</div>

Oh yeah, a compare and contrast with Query would be helpful for understanding why this exists and what ppl should look for when evaluating it

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 31, 2019, 3:19am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/12 "2019-01-31T03:19:09Z")

</div>

Hmm, well it’s similar, but different in a couple of ways:

1. Efficient usage with native `missing`
2. No reliance on inference (I’m not sure if it still does, but Query did used to rely on inference)
3. Much simpler interface and code (two very simple macros, and only two new iterators. The rest of the iterators are all taken from Base.Iterators).
4. Added flexibility from relying on an explicit chaining macro
5. Not tested against a huge variety of data sources and sinks. I think? it should be theoretically possible to support them.
6. Huge performance improvements for presorted data with grouping and joining

---

<div class="post-metadata">

**Author:** ![xiaodai](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/xiaodai/32/15937_2.png) [@xiaodai](https://discourse.julialang.org/u/xiaodai)\
**Post date:** [January 31, 2019, 3:30am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/13 "2019-01-31T03:30:43Z")

</div>

> [@bramtayl](#):
>
> Huge performance improvements for presorted data with grouping and joining

This is interesting!

---

<div class="post-metadata">

**Author:** ![mkborregaard](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mkborregaard/32/556_2.png) [@mkborregaard](https://discourse.julialang.org/u/mkborregaard)\
**Post date:** [January 31, 2019, 9:20am UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/14 "2019-01-31T09:20:06Z")

</div>

Nice. I notice when following the tutorial that it’s a lot more verbose than dplyr. Is that by design, or do you plan to simplify the syntax?

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 31, 2019, 4:48pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/15 "2019-01-31T16:48:53Z")

</div>

I have some simplifications planned for groups and joins.

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 31, 2019, 4:53pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/16 "2019-01-31T16:53:23Z")

</div>

Not much I can do about the `@_ _.column` syntax, I’m afraid. Julia doesn’t have lazy evaluation like R does.

---

<div class="post-metadata">

**Author:** ![ValdarT](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/valdart/32/24146_2.png) [@ValdarT](https://discourse.julialang.org/u/ValdarT)\
**Post date:** [January 31, 2019, 6:08pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/17 "2019-01-31T18:08:06Z")

</div>

I don’t want to derail the discussion but personally I’m a bit sad that the [StructuredQueries](https://julialang.org/blog/2016/10/StructuredQueries) project hasn’t continued. That approach makes a lot of sense because it can provide both a nice syntax as well as support for “external backends” such as databases or Spark. I think the combination of these two aspects is extremely powerful and the main reasons behind the success of `dplyr` and `SQL`.  
The point here being that perhaps it’s worthwhile to focus more on the syntax and abstracting away the “backend” even at the cost of added complexity in implementation. And just in case: this is not meant as a criticism on any level, just a comment regarding

> [@bramtayl](#):
>
> in comparison to other querying packages

---

<div class="post-metadata">

**Author:** ![bramtayl](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/bramtayl/32/3614_2.png) [@bramtayl](https://discourse.julialang.org/u/bramtayl)\
**Post date:** [January 31, 2019, 6:27pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/18 "2019-01-31T18:27:38Z")

</div>

So just to give a flavor of the kind of syntax improvements I could make, consider:

```julia
julia> dest_tailnum =
          @> flights |>
          rows |>
          order(_, select(:dest, :tailnum)) |>
          By(_, select(:dest, :tailnum)) |>
          Group |>
          over(_, @_ transform(_.first,
                    flights = length(_.second)
          )) |>
          columns(_, :dest, :tailnum, :flights)

```

I could provide some convenience functions, for example:

```julia
julia> dest_tailnum =
          @> flights |>
          group_by(_, :dest, :tailnum) |>
          summarize(_, flights = @_ length(_.flights)) |>
          ungroup

```

Which is pretty darn close to the dplyr syntax. Of course, there’s a loss in terms of flexibility and clarify (IMHO), but sounds like something people would really like?

---

<div class="post-metadata">

**Author:** ![davidanthoff](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/davidanthoff/32/223493_2.png) [@davidanthoff](https://discourse.julialang.org/u/davidanthoff)\
**Post date:** [January 31, 2019, 7:30pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/19 "2019-01-31T19:30:48Z")

</div>

> [@ValdarT](#):
>
> That approach makes a lot of sense because it can provide both a nice syntax as well as support for “external backends” such as databases or Spark. I think the combination of these two aspects is extremely powerful and the main reasons behind the success of `dplyr` and `SQL` .

[Query.jl](https://github.com/queryverse/Query.jl)s design is all setup to support that scenario, and at one point I had an example of a very simple translation to SQL. That was only a proof of concept, and would require a fair bit of work to make actually usable, but in terms of architecture everything is in place to support that kind of scenario.

---

<div class="post-metadata">

**Author:** ![affans](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/affans/32/11911_2.png) [@affans](https://discourse.julialang.org/u/affans)\
**Post date:** [January 31, 2019, 7:47pm UTC](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262/20 "2019-01-31T19:47:37Z")

</div>

So what does one use? `query.jl` or this package? Is there a way to merge them? I.e. take the best features of both ?😃

[Next page](https://discourse.julialang.org/t/ann-lightquery-tutorial/20262.md?page=2)
