CausalGraphs.jl: from measurement models to causal graphs

Hello Julia community,

A week ago, I mentioned some work around measurement models and uncertainty analysis in the "What projects have you been working on?" section of Julia Discourse.

Since then, that work has evolved into a more general abstraction for representing cause-effect relationships, which led me to start (a tiny package) named CausalGraphs.jl.

image

The initial motivation came from measurement models, uncertainty analysis, and Ishikawa (fishbone) diagrams. While working on this, I realized that the causal/dependency structure itself was worth representing as a reusable data model, rather than treating the Ishikawa diagram as the primary object.

So I extracted this part into a dedicated Julia package.

CausalGraphs.jl is intended for building, manipulating, analyzing, and visualizing cause-effect knowledge graphs.

The central idea is to separate:

  • Knowledge: causes, effects, measurements and their relationships
  • Views: Ishikawa/fishbone diagrams, dependency graphs and other visualizations
  • Models: mathematical and measurement models, uncertainty propagation, fault trees, etc.

This separation is important because an Ishikawa diagram is only one possible representation of the underlying structure.

Connecting causal structure and uncertainty

This also connects to another project I’ve been working on: SymbolicUncertainties.jl.

SymbolicUncertainties.jl is not public yet, but I’m making the documentation available while the design evolves.

The two projects are intended to address complementary aspects of the same problem:

              CausalGraphs.jl
                     |
                     | What depends on what?
                     | What are the contributing factors?
                     v
             Mathematical model
                     |
                     | How is the quantity computed?
                     v
          SymbolicUncertainties.jl
                     |
                     | How does uncertainty propagate?
                     | Which inputs dominate the result?
                     v
              Quantitative analysis

For example, an engineering measurement might have a causal structure involving temperature, calibration, sensor characteristics, environmental effects, acquisition noise, etc.

The graph describes how these quantities are related.

A mathematical model then describes how the measured quantity is calculated.

Finally, symbolic uncertainty analysis can describe how uncertainty propagates through that model.

This is one of the directions I’m exploring in relation to the GUM approach to measurement uncertainty.

Why make the graph a first-class object?

One of the things I’m interested in is whether this abstraction could be useful beyond measurement uncertainty.

The same causal/dependency graph could potentially be used for:

  • measurement and uncertainty analysis
  • root-cause analysis
  • reliability engineering
  • machine and system safety
  • fault-tree analysis
  • quality engineering
  • scientific and engineering models
  • sensitivity analysis

For example, an Ishikawa diagram could simply be one view of the graph, while other tools could operate directly on the graph structure.

This also potentially makes it possible to connect qualitative analysis with quantitative models without forcing either side to become the other.

Julia ecosystem

I’d particularly appreciate feedback on how this should fit into the existing Julia ecosystem.

In particular, I’m interested in the relationship with packages such as:

  • Graphs.jl
  • Makie
  • uncertainty quantification tools
  • symbolic modeling
  • reliability and safety analysis

Some questions I’m currently exploring:

  1. Is this separation between causal knowledge, views, and mathematical models a useful abstraction?

  2. Should CausalGraphs.jl build more explicitly on Graphs.jl?

  3. What is the best way to associate a causal/dependency graph with a symbolic mathematical model?

  4. Could this provide a useful bridge between qualitative causal analysis and quantitative uncertainty analysis?

  5. Are there existing Julia projects that already address parts of this problem that I should integrate with?

The project is still at an early stage (and is not registered in Julia General Registry), so I’m particularly interested in architecture and API feedback.

CausalGraphs.jl: GitHub - s-celles/CausalGraphs.jl: Julia package for building, manipulating, and analyzing cause-effect knowledge graphs · GitHub
Documentation: Home · CausalGraphs.jl
SymbolicUncertainties.jl: Home · SymbolicUncertainties.jl
Some rendering examples:

Graphs.bfs_tree(g, s[; dir=:out])

Provide a breadth-first traversal of the graph g starting with source vertex s, and return a directed acyclic graph of vertices in the order they were discovered. If dir is specified, use the corresponding edge direction (:in and :out are acceptable values).

and

DirectedAcyclicGraphs

address aspects of this. Causal inference depends on controlling for confounding and conflcting variable that muddy the waters between causes and effects (there seem always to be multiples of each), See Judiah Pearl’s The Book of Why if you haven’t already.

The separation of causal knowledge, views and mathematical models is core to creating an ontology, to use the current buzzword. It’s not enough to connect the dots (vertices), it’s also necessary to weight the edges, and descending levels of abstraction are needed to make that feasible to overcome the 7±2 register stack that wetware works with.

Take a look at GitHub - Shmuma/Dagitty.jl: Graphical Analysis of Structural Causal Models · GitHub and https://www.dagitty.net/.

Thanks. I noticed I did a mistake while creating this announcement.

There is ever a CausalGraphs.jl package in Julia General Registry

from @xiangao Home · CausalGraphs.jl

Sorry about that

There are also at least two packages for the data analysis side of things:

(and we are working on briding the two)

There is also the package I have been working on for a couple of months. It’s for working with causal graphs, so of course DAGs like we discuss here, but also equivalence classes (PDAGs/MPDAGs/CPDAGs), with latent variables (ADMGs), Ancestral Graphs (AGs and MAGs), and equivalence classes of MAGs (PAGs).

It also have an extensive plotting support of causal graphs. I’ll make a separate post soon when I release 1.0.

https://bjarkehautop.github.io/CausalStructures.jl/stable/

(Feel free to reach out here or on slack)

As CausalInference.jl takes the name of the whole field in Julia, it would be awesome if that package became the high-level interface for the whole field. Maybe there will be an organization at some point. (I was confused by the content at first because causal discovery algorithms are barely mentioned in my field.)

I think the name CausalInference.jl is overselling a bit. The whole field is substantially more generic than the functionality the package covers. I think CausalGraphInference.jl is a more appropriate name.

But yes an organisation may be helpful since there is some commynity interest.

I think the interesting question is not only whether we should create a JuliaCausal organization, but also how the different packages should fit together.

For example, I could see the following roles:

  • CausalGraphs.jl for representing causal, dependency and measurement knowledge

  • CausalStructures.jl for formal structures such as DAGs, ADMGs, CPDAGs, PAGs, etc.

  • Associations.jl for statistical associations and causal discovery

  • CausalInference.jl for identification, adjustment and estimation, potentially providing a higher-level interface

There are several directions that could be developed from there. In particular, I think interoperability between the packages is more important than having a single package do everything.

For this CausalGraphs.jl, this could mean richer nodes and edges with things like evidence, confidence, provenance, time relationships and levels of abstraction. The same underlying graph could then be used to produce different views such as causal graphs, dependency graphs, Ishikawa diagrams or fault trees.

Another interesting direction would be connecting these graphs to mathematical models, for example through ModelingToolkit, and eventually to quantitative analysis such as uncertainty propagation and sensitivity analysis.

If we can agree on these kinds of interfaces and boundaries, the organization itself should become much easier to define.

This might be the wrong place to ask, but I thought the right people might be here.
I’m developing ScientificFitting.jl and see a potentially useful connection here. ScientificFitting.jl basically does the fits of some model function to some real world data and makes it easy to understand the uncertainty structure, parameter correlations, goodness of fit etc.
But from there one might be interested in the causal relations between some measurement and some other data, other fits, etc.
So I think that it would be useful to be able to simply parse the results to some of the packages here. Something like:

Measured data → fit → correlated parameter estimates → uncertainty propagation in a larger measurement model.

The important part would be preserving parameter correlations, rather than exporting independent ± values. Fit diagnostics and profiles would remain available alongside that representation. In other words, ScientificFitting.jl might become one very useful input for this sorts of causal graph / inference analysis.
But now there are many different packages in this area and no real common interface and I am kind of confused how this would look like in practice.

  • Is this even something you are interested in?
  • Is there any preferred package / common API for which I could develop?

And I agree with some of the earlier comments that it would be great to have an organization for this and also a common interface.
Any ideas on this?