# Designing Custom Syntax for @formula

**URL:** <https://discourse.julialang.org/t/designing-custom-syntax-for-formula/80450>\
**Category:** Statistics\
**Tags:** question, statsmodels\
**Created:** [May 3, 2022, 10:09pm UTC](https://discourse.julialang.org/t/designing-custom-syntax-for-formula/80450 "2022-05-03T22:09:38Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ross\_Boylan](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ross_boylan/32/9210_2.png) [@Ross\_Boylan](https://discourse.julialang.org/u/Ross_Boylan)\
**Post date:** [May 3, 2022, 10:09pm UTC](https://discourse.julialang.org/t/designing-custom-syntax-for-formula/80450/1 "2022-05-03T22:09:38Z")

</div>

I’m trying to understand the differences between the relatively compact implementation of a [custom interpretation of `^`](https://github.com/kleinschmidt/RegressionFormulae.jl/blob/main/src/power.jl) for `@formula` and the general advice on extending `@formula`, in particular the example [here](https://juliastats.org/StatsModels.jl/stable/internals/#An-example-of-custom-syntax:-poly) of how to implement a custom interpretation for `poly()`.

The main source of the difference is that the latter creates a custom `Term` type, `PolyTerm` to hold the expression, and then implements a bunch of methods to deal with that type. In contrast, the code for `^` expands the terms out immediately in `apply_schema` without any reference to new term types.

Why the differences, and which is a better model to use?

Perhaps without something like a `PowerToTerm` for `^` it would be harder to construct formulae programmatically?

Thanks.

---

<div class="post-metadata">

**Author:** ![dave.f.kleinschmidt](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/dave.f.kleinschmidt/32/55_2.png) [@dave.f.kleinschmidt](https://discourse.julialang.org/u/dave.f.kleinschmidt)\
**Post date:** [May 13, 2022, 3:23pm UTC](https://discourse.julialang.org/t/designing-custom-syntax-for-formula/80450/2 "2022-05-13T15:23:14Z")

</div>

I actually go back and forth on this issue. Originally in StatsModels.jl, ALL of the special syntax stuff happened at parse time, inside the macro. So the formula that comes out at the end of `a * b` has `a + b + a&b` and no memory of `a * b`. Recently we’ve started to move more of that into run time by adding methods for things like `Base.:*(a::Term, b::Term) = a + b + a & b`, rather than having a transformation that works on the `Expr` that hte macro sees. IIRC all that stuff is languishing in [https://github.com/JuliaStats/StatsModels.jl/pull/183](https://github.com/JuliaStats/StatsModels.jl/pull/183) and I’ve had some second thoughts in the intervening time. Adding all those methods puts an even bigger burden on the compiler, but using a very differnet approach would require even more dramatic internal (and possibly external) changes (e.g., could be hard to have stuff like `term(:a) * term(:b)` work without defining those methods).

So, all of which is to say, it’s a design decision, and there’s no obviously correct choice 🙂 I’d say that the best starting point is PROBABLY to start with a `PolyTerm`-like approach, rather than the `^` approach. It’s a lot simpler to get the bookkeeping right if you have a 1-1 match between the input `FunctionTerm` and the output terms. It’s possible to handle 1-to-many transforms but it can get fiddly (see some of the stuff around `/` for instance, or how `/` is handled on the RHS of random effects in MixedModels.jl)
