
UNSW Sydney | STREAM

We are targeting the quantity \(p(f \mid \mathcal{C})\).

| prior | observation model | what you get | |
|---|---|---|---|
| OI / 3D-Var | Gaussian, static \(\mathbf{B}\) | linear \(H\), Gaussian | a mean and a covariance |
| 4D-Var | Gaussian, model as strong constraint | linearised \(H\), Gaussian | the mode, via the adjoint |
| EnKF | an ensemble, Gaussian in the update | linear update | a shifted ensemble |
| this talk | your model \(f\) | almost anything | draws from the true posterior |

Diffusion modelling simply asks “how do I sample from an inconvenient distribution?”

Anderson (1982), Reverse-time diffusion equation models, Stochastic Processes and their Applications 12(3):313–326.


Price, Ilan, et al. (2025), Probabilistic weather forecasting with machine learning., Nature 637(8044): 84-90.


The marginal \(p_t\) stays Gaussian, and we know \[ p_t = \mathcal{N}\big(\boldsymbol{b}(t),\, \mathbf{A}(t)\big), \qquad \boldsymbol{b}(t) = \alpha(t)\,\mathbf{m}_*, \qquad \mathbf{A}(t) = \alpha^2(t)\mathbf{K}_{**} + (1-\alpha^2(t))\mathbf{I}. \]
So the score is closed form, and no network is required:
\[ \nabla_{\boldsymbol{f}}\log p_t(\boldsymbol{f}_t) = -\mathbf{A}(t)^{-1}\big(\boldsymbol{f}_t - \boldsymbol{b}(t)\big). \]
Drop it into the reverse-time SDE and integrate from \(t=1\) to \(t=0\):
\[ \mathrm{d}\boldsymbol{f}_t = \beta(t)\left[\tfrac{1}{2}\boldsymbol{f}_t - \nabla_{\boldsymbol{f}}\log p_t(\boldsymbol{f}_t)\right]\mathrm{d}t + \sqrt{\beta(t)}\,\mathrm{d}\bar{\boldsymbol{W}}_t. \]
A GP/OI is a very boring cat with a score we can write down.
“The most efficient thing for a computer to do is nothing” - Prof. Philipp Hennig
Whiten: \(\hat{\boldsymbol{f}}_0 = \mathbf{K}_{**}^{-1/2}(\boldsymbol{f}_0 - \mathbf{m}_*)\), the dynamics vanish entirely … the linear drift term is zero.
Rosenblatt (1952), Remarks on a multivariate transformation, The Annals of Mathematical Statistics 23(3):470–472.
Every one is \(\boldsymbol{f}_0 = T_\vartheta(\boldsymbol{\xi}_0)\) for a tractable generator \(T_\vartheta\) and iid-normal latent \(\boldsymbol{\xi}_0\):
\[ \nabla \log p_t(\boldsymbol{f}_t \mid \mathcal{C}) = \nabla \log p_t(\boldsymbol{f}_t) + \nabla \log p_t(\mathcal{C} \mid \boldsymbol{f}_t). \]
Optimal interpolation, conditioned on the observations alone

Conditioned on the observations and on the non-linear damped pendulum equation

We work in latent space. Target the pulled-back posterior: \[ \pi(\boldsymbol{\xi}_0 \mid \mathcal{C}) \propto p\big(\mathcal{C} \mid T_\vartheta(\boldsymbol{\xi}_0)\big) \rho(\boldsymbol{\xi}_0). \]
Along the latent flow, the score decomposes: \[ \nabla_{\boldsymbol{\xi}} \log p_t(\boldsymbol{\xi}_t \mid \mathcal{C}) = \underbrace{\nabla_{\boldsymbol{\xi}} \log p_t(\boldsymbol{\xi}_t)}_{=\,-\boldsymbol{\xi}_t\ \text{(exact)}} + \underbrace{\nabla_{\boldsymbol{\xi}} \log p_t(\mathcal{C} \mid \boldsymbol{\xi}_t)}_{\text{guidance}}. \]
Run the same SDE in \(\boldsymbol{\xi}\)-space, add the guidance term, then push through \(T_\vartheta\).
The latent prior is exact at every \(t\). Guidance is the only thing to approximate.
The increment is the pull towards \(\mathcal{C}\), averaged over where the current draw could still end up: \[ \nabla_{\boldsymbol{\xi}}\log p_t(\mathcal{C}\mid \boldsymbol{\xi}_t) \approx \alpha(t) \sum_{i=1}^{S} \bar{w}^{(i)}\, \nabla_{\boldsymbol{\xi}}\log p\big(\mathcal{C}\mid T_\vartheta(\boldsymbol{\xi}_0^{(i)})\big), \qquad \bar w^{(i)} \propto p\big(\mathcal{C}\mid T_\vartheta(\boldsymbol{\xi}_0^{(i)})\big). \]
(There are known weight-degeneracy issues in importance sampling that do not appear to matter here, for a handful of esoteric reasons.)
| Error source | Image diffusion | Simulator-based | |
|---|---|---|---|
| 1 | Marginal score \(\nabla\log p_t\) | uncontrolled | exact |
| 2 | Guidance approximation | uncontrolled | explicit; \(\mathcal{O}(S^{-1})\) bias |
| 3 | Discretisation | controllable | solver order \(h^q\), same bound |
The conditioning mechanism does not care where the score comes from. Closed form from your solver, or learned from data.
\[ \mathrm{d}T = -\kappa(S)\,\bigl(T - T^{*}(S)\bigr)\,\mathrm{d}t + \sigma_T\,\mathrm{d}W_T, \qquad \mathrm{d}S = 1/C_S \bigl(P(t) - E(T,S)\bigr)\,\mathrm{d}t \]
\(\kappa(S)\) relaxation rate · \(T^{*}(S)\) equilibrium temperature · \(C_S\) soil water storage capacity · \(P(t)\) precipitation · \(E(T,S)\) evaporative water loss
An insultingly simplified version of Brubaker & Entekhabi (1996).
NN + Diffusion based inversion of methane flux

Computational scaling vs true-model inversion




We’ve done a lot of theory this year:
Now, can we actually DO anything?!

Slides and code: astfalckl.github.io/presentations
STREAM: unsw.edu.au/science/our-schools/maths/our-research/stream
l.astfalck@unsw.edu.au
Lachlan Astfalck | UNSW Spatio-Temporal Research for Environmental Analysis and Modelling