[ANN] MissingPatterns.jl — visualize missing data in the terminal

Hi everyone,

I’d like to announce MissingPatterns.jl, a small toolkit for exploring missing data patterns directly in the terminal.

The package started from a gap I kept running into: in Python I’d reach for missingno for a quick look at missingness before anything else, and in R for mice::md.pattern() or naniar. Coming to Julia I looked for an equivalent and didn’t find one, the usual answer was to compute the masks by hand and plot them yourself. So I wrote the tool I wanted to have, with one difference from missingno: since these are diagnostics you run early and often, I wanted them to render in the REPL without pulling in a plotting stack.

It works with any Tables.jl-compatible source (DataFrame, CSV.File, NamedTuple of vectors, …) and has zero plotting dependencies, everything is rendered with Unicode/ANSI, so it also works over SSH and inside IDE/Jupyter output cells.

using Pkg; Pkg.add("MissingPatterns")

using MissingPatterns, DataFrames

df = DataFrame(
    A = [1, missing, 3, missing, 5, 6, 7, missing],
    B = [missing, 2, 3, missing, 5, 6, 7, missing],
    C = [1, 2, 3, 4, missing, 6, 7, 8],
)

missingpatterns(df)
┏━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┓
┃    A    ┃    B    ┃    C    ┃    n    ┃    %    ┃  freq   ┃
┣━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃    3    ┃  37.5%  ┃ ███████ ┃
┃  █████  ┃  █████  ┃  ░░░░░  ┃    2    ┃  25.0%  ┃ █████   ┃
┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃    1    ┃  12.5%  ┃ ██      ┃
┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃    1    ┃  12.5%  ┃ ██      ┃
┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃    1    ┃  12.5%  ┃ ██      ┃
┗━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┛

 5 unique patterns across 8 rows

A worked example

Toy tables don’t show much, so here is a full pass over something closer to real data: one year of daily readings from a weather station, with the kind of missingness these files actually have, a probe that went offline for three weeks, a sensor that was down for a whole month, and background dropouts everywhere.

using MissingPatterns, DataFrames, Dates, Random, Statistics
Random.seed!(42)

dates = Date(2024,1,1):Day(1):Date(2024,12,31)
n = length(dates)

F = Vector{Union{Missing,Float64}}

df = DataFrame(
    date   = collect(dates),
    temp   = F(round.(15 .+ 10 .* sin.(2pi .* (1:n) ./ 365) .+ randn(n), digits=1)),
    humid  = F(round.(60 .+ 15 .* randn(n), digits=1)),
    wind   = F(round.(abs.(3 .+ 2 .* randn(n)), digits=1)),
    precip = F(round.(abs.(randn(n)) .* 5, digits=1)),
    pm25   = F(round.(20 .+ 8 .* randn(n), digits=1)),
)

# the temperature/humidity probe was offline for three weeks in March
df[Date(2024,3,4) .<= df.date .<= Date(2024,3,22), [:temp, :humid]] .= missing

# the PM2.5 sensor was down for the whole of August
df[month.(df.date) .== 8, :pm25] .= missing

# plus sporadic dropouts
for (col, k) in ((:pm25, 25), (:humid, 18), (:temp, 10), (:wind, 6))
    df[randperm(n)[1:k], col] .= missing
end

missingsummary — how much, and where along the rows

The first question is always “how bad is it, and is it spread out or clumped?”. Each column gets a count, a percentage, and a sparkline over the row axis (20 bins by default), so a block outage and a uniform trickle stop looking like the same number.

missingsummary(df)
 column  type       missing        %  distribution
 pm25    Float64         54   14.75%   ▁▂▁▁   ▁▁▁▇█▁▁▁▂▁▁▂
 humid   Float64         36    9.84%   ▁▁▆▃▁  ▁▁▁▁▁ ▁▁▁▁
 temp    Float64         28    7.65%  ▁  ▆▃ ▁ ▁▁ ▁▁ ▁
 wind    Float64          6    1.64%  ▁      ▁   ▁ ▁
 date    Date             0    0.00%
 precip  Float64          0    0.00%
 124 missing of 2196 cells (5.65%) across 6 columns ┊ bins of 19 rows

pm25 has one tall spike (the August outage) sitting on a low background; temp and humid share a bump in the same place, early in the year. A bin with even a single missing value always renders at least , so isolated holes never vanish into a blank block.

plotmissing — the heatmap, optionally on a calendar axis

plotmissing(df) gives the classic missingno-style matrix, compressed with a block gradient when the table is taller than the terminal. What I use most, though, is by=:date, period=:month: the vertical axis stops being arbitrary row ranges and becomes honest calendar time.

plotmissing(df; by=:date, period=:month, name_width=6)
┏━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┓
┃       ┃    0%   ┃    8%   ┃   10%   ┃    2%   ┃    0%   ┃   15%   ┃
┣━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃  row  ┃  date   ┃  temp   ┃  humid  ┃  wind   ┃ precip  ┃  pm25   ┃
┣━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃2024-01┃  ░░░░░  ┃  ·····  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃
┃2024-02┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ▒▒▒▒▒  ┃
┃2024-03┃  ░░░░░  ┃  █████  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃
┃2024-04┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-05┃  ░░░░░  ┃  ·····  ┃  ░░░░░  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-06┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-07┃  ░░░░░  ┃  ·····  ┃  ·····  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-08┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃  ░░░░░  ┃  █████  ┃
┃2024-09┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-10┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-11┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃
┃2024-12┃  ░░░░░  ┃  ░░░░░  ┃  ·····  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃
┗━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┛

MissingPatterns.Analysis: 366 × 6 DataFrame
 Grouping:    by date (month) → 12×6 cells
 Missing (count):             124               ┊ Missing (%):           5.65%
 Present (count):            2072               ┊ Present (%):          94.35%
 Progress Bar:     [██████████████████████████████████████████████████████████]

Each cell is shaded by the missing fraction of its block, from fully present () through ·, and up to fully missing (), so the March probe outage and the August sensor failure are readable at a glance and dated. period also takes :year, :quarter, :week and :day, and by alone (no period) groups by the raw values of any sortable column — by=:station, by=:region — which turns the same plot into “is this missingness concentrated in one group?”.

In a color terminal, layout=:compact packs two grid rows per output line with half-blocks (use color=:always in VS Code or Jupyter, where stdout isn’t a TTY but ANSI still renders).

missingpatterns — which columns go missing together

The equivalent of R’s mice::md.pattern(): unique missingness signatures, sorted by frequency.

missingpatterns(df; name_width=6, max_patterns=8)
┏━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┓
┃  date   ┃  temp   ┃  humid  ┃  wind   ┃ precip  ┃  pm25   ┃    n    ┃    %    ┃  freq   ┃
┣━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃   268   ┃  73.2%  ┃ ███████ ┃
┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃   49    ┃  13.4%  ┃ █       ┃
┃  ░░░░░  ┃  █████  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃   19    ┃  5.2%   ┃         ┃
┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃   14    ┃  3.8%   ┃         ┃
┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃    6    ┃  1.6%   ┃         ┃
┃  ░░░░░  ┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃    4    ┃  1.1%   ┃         ┃
┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃  █████  ┃    2    ┃  0.5%   ┃         ┃
┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  █████  ┃  ░░░░░  ┃  ░░░░░  ┃    1    ┃  0.3%   ┃         ┃
┗━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┛

 11 unique patterns across 366 rows  ┊ showing top 8 of 11

73.2% of the days are complete; the third row is the probe outage, the one signature where temp and humid disappear jointly — that’s the one that matters for a complete-case analysis, and the one an independent per-column imputation would get wrong. min_pct=1.0 collapses the long tail of one-off signatures instead of capping by count.

missingcooccurrence — pairwise phi / Jaccard between masks

The patterns table says which combinations exist; this one summarises the pairwise dependence between missingness masks, which is what you actually cite when arguing against MCAR.

missingcooccurrence(df; name_width=6)      # method=:jaccard for the overlap index
┏━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┓
┃    ϕ    ┃  date   ┃  temp   ┃  humid  ┃  wind   ┃ precip  ┃  pm25   ┃
┣━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃  date   ┃    —    ┃    ·    ┃    ·    ┃    ·    ┃    ·    ┃    ·    ┃
┃  temp   ┃    ·    ┃    —    ┃  0.60   ┃  0.04   ┃    ·    ┃  -0.06  ┃
┃  humid  ┃    ·    ┃  0.60   ┃    —    ┃  -0.04  ┃    ·    ┃  -0.06  ┃
┃  wind   ┃    ·    ┃  0.04   ┃  -0.04  ┃    —    ┃    ·    ┃  0.01   ┃
┃ precip  ┃    ·    ┃    ·    ┃    ·    ┃    ·    ┃    —    ┃    ·    ┃
┃  pm25   ┃    ·    ┃  -0.06  ┃  -0.06  ┃  0.01   ┃    ·    ┃    —    ┃
┗━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┛
 pairwise ϕ of missingness masks ┊ n = 366 rows

ϕ(temp, humid) = 0.60 against ~0 everywhere else: the two probes fail as one device, everything else fails independently. Columns with no missing values are · (degenerate, no coefficient defined) and the diagonal is . With color on, cell intensity scales with the absolute value.

plotmissingdiff — auditing an imputation step

Same table before and after a transformation, cell by cell. Here I fill the three meteorological columns with their column means and deliberately leave pm25 alone:

imputed = copy(df)
for c in (:temp, :humid, :wind)
    imputed[!, c] = coalesce.(imputed[!, c], round(mean(skipmissing(imputed[!, c])), digits=1))
end

plotmissingdiff(df, imputed; name_width=6, target_lines=14)
┏━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━┓
┃ da… +0% ┃ te… -8% ┃ h… -10% ┃ wi… -2% ┃ pr… +0% ┃ pm… +0% ┃
┣━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━╋━━━━━━━━━┫
┃  ·····  ┃  -----  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  ·····  ┃  -----  ┃  ·····  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  -----  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  -----  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  -----  ┃  -----  ┃  -----  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  ·····  ┃  -----  ┃  ·····  ┃  ·····  ┃  ·····  ┃
┃  ·····  ┃  ·····  ┃  -----  ┃  ·····  ┃  ·····  ┃  ·····  ┃
┗━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┻━━━━━━━━━┛
 Δ missing: 5.65% → 2.46% ┊ resolved 70 ┊ introduced 0

- marks blocks with fewer missing values than before, + blocks with more, · unchanged; the per-column header carries the signed delta and the footer the exact cell counts, computed by a row-aligned pass rather than from the block averages. In color, resolved blocks are green and introduced ones red. The point of the plot is pm25 +0% and introduced 0: the step did what it claimed, and nothing else, which is exactly the assertion that is easy to make and hard to check after a pipeline touches a dataset.

missinghtml — the same heatmap, portable

For reports and notebook exports, the same compression engine and color ramp render to a self-contained <div>, no external CSS or JS, every cell carrying a tooltip with its row range and exact percentage. It affords a much bigger grid than a terminal (defaults: 200x60 blocks).

html = missinghtml(df; title = "Station 08A - 2024")   # returns a String
missinghtml("station08a.html", df; title = "Station 08A - 2024", max_rows = 180)

Wrapping up

Full docs: Home · MissingPatterns.jl

Feedback, bug reports and ideas are very welcome — especially on the compressed rendering for large tables, and on which diagnostics would be worth adding next.

Nicely Done! This is cool!