Compare predictions#

Use pointwise log likelihoods to evaluate how a fitted model predicts held-out observations with Pareto-smoothed importance sampling leave-one-out cross-validation (PSIS-LOO). Start with checked chains; LOO does not diagnose MCMC convergence.

import jax.numpy as jnp
import numpy as np
import tensorflow_probability.substrates.jax.bijectors as tfb
import tensorflow_probability.substrates.jax.distributions as tfd

import liesel.goose as gs
import liesel.model as lsl

Prepare the example#

These examples require the regression model and sampling results from Sample your first posterior.

Compute log likelihoods#

For model and results from Sample your first posterior:

samples = results.get_posterior_samples()
log_lik = lsl.log_prob_pointwise({"y": model.vars["y"]}, samples)
log_lik_y = log_lik["y_log_prob"]
loo_result = gs.loo(log_lik_y)
log_lik_y.shape
(4, 1000, 500)
loo_result
Computed from 4000 posterior samples and 500 observations log-likelihood matrix.

         Estimate       SE
elpd_loo  -721.24    16.56
p_loo        3.09        -
------

Pareto k diagnostic values:
                         Count   Pct.
(-Inf, 0.70]   (good)      500  100.0%
   (0.70, 1]   (bad)         0    0.0%
    (1, Inf)   (very bad)    0    0.0%

The shape is (4, 1000, 500): chains, draws, observations. Select the response likelihood, not the total log posterior or its priors. The distribution must retain per-observation log probabilities. See log_prob_pointwise() for the exact rules.

Read the result#

elpd_loo estimates expected log predictive density under leave-one-out prediction; larger is better on the default log scale. se describes uncertainty in that estimate. Inspect the reported Pareto-k diagnostics and warnings: unreliable importance weights can make the approximation unsuitable. Here all 500 Pareto-k values are in the reported good range. p_loo is the estimated effective number of parameters (about 3.07 here); it gives context for model complexity and is not the predictive score to maximize. The ArviZ PSIS-LOO documentation explains the result fields and diagnostics.

The optional samples argument to loo() is retained for compatibility but ignored. By default, relative MCMC efficiency is estimated from likelihood values. Use gs.loo(log_lik_y) for this workflow.

Compare like with like#

Fit each candidate model and evaluate the same observations, in the same order and on the same likelihood scale. Compare their predictive scores with uncertainty, not just a ranking. The standard error for an individual score is not the standard error of a paired difference between two models.

The observation axis defines what is left out. Leaving out one response is a different prediction task from leaving out a group or a future time interval. Do not flatten dependent likelihood factors into purported independent observations without checking that the resulting task matches your analysis.