Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Proposal / feedback: topology-aware posterior predictive and simulator summaries

Дата публикации: 04-10-2026 08:17:59

Proposal / feedback: topology-aware posterior predictive and simulator summaries
Hi everyone,
I recently contributed the topological posterior predictive checks example in pymc-examples (#881).
During review, an important point came up: in the seasonal toy example I used, the missing structure could also be detected using simpler diagnostics such as autocorrelation or the Fourier spectrum.
It is a legitimate criticism, so before proposing anything beyond the example notebook I tried to answer a more specific question:
Is there a useful class of Bayesian model-checking or simulation-inference problems where persistent homology contains information that is genuinely unavailable to Fourier-spectrum and autocorrelation diagnostics?
I have now run a sequence of synthetic benchmarks, including several negative ones, and I think there might be a clean use case worth discussing.
The construction
I used two spatial point-pattern mechanisms based on an exact homometric pair.
The two latent patterns have the same displacement multiset. Consequently, they have the same complete two-point autocorrelation:
C_A(h)=C_B(h),
and therefore the same Fourier power at every frequency:
|\widehat{\mu_A}(k)|^2
=
|\widehat{\mu_B}(k)|^2.
So this is stronger than saying that their spectra or autocorrelations happen to be similar: they are identical at the level of complete second-order structure.
However, the way those pairwise relationships are assembled into higher-order spatial structure differs.
In particular, the two mechanisms have different H_1 persistent-homology signatures.
Persistent homology can therefore distinguish the mechanisms even though second-order statistics cannot.
What happens under observation noise?
I then deliberately degraded the observations using:
coordinate-localization noise;
random point dropout;
finite spatial resolution / quantization.
At that point exact homometry is broken, so Fourier and autocorrelation are allowed to become informative.
The conventional baseline explicitly contained Fourier power, a direct 2D empirical displacement/autocorrelation histogram, pairwise-distance summaries, and covariance / ordinary geometry.
Persistent-homology features were then added on top of exactly that same baseline.
The topological advantage survived substantial observation error.
Benchmark
Conventional
+ persistent homology
Exact homometry, untouched-test AUC
0.500
1.000
Moderate+ corruption, replicated median AUC
0.473
0.847
Median AUC lift
—
+0.359
90% replication interval for AUC lift
—
[0.302, 0.451]
Median log-loss reduction
—
29.7%
Independent test sets clearing both practical lift criteria
—
20 / 20
The moderate+ corruption regime used:
coordinate jitter SD = 0.50
15% point dropout
spatial quantization resolution = 0.60
So the effect does not disappear as soon as the exact mathematical construction is perturbed.
Does this actually help Bayesian inference?
I also wanted to avoid stopping at classification.
I built a simulator in which each observed spatial field is generated from one of two mechanisms:
M_i \sim \mathrm{Bernoulli}(w),
with
w \sim \mathrm{Beta}(1,1),
where w is the unknown population fraction generated by mechanism M_1.
The latent mechanism labels M_i are not observed.
A calibrated classifier trained with equal model probabilities estimates
q(y)
\approx
P(M=1\mid y).
Because the classifier-training prior is balanced,
P(M=0)=P(M=1)=\frac12,
its odds estimate the likelihood ratio:
r(y)
\approx
\frac{q(y)}{1-q(y)}
\approx
\frac{p(y\mid M=1)}
{p(y\mid M=0)}.
For the mixture model,
p(y_i\mid w)
=
(1-w)p(y_i\mid M=0)
+
w\,p(y_i\mid M=1).
Factoring out p(y_i\mid M=0) gives
p(y_i\mid w)
\propto
(1-w)+w\,r(y_i).
Therefore the posterior is, up to a multiplicative constant,
p(w\mid y_{1:n})
\propto
p(w)
\prod_{i=1}^{n}
\left[
(1-w)+w\,r(y_i)
\right].
I implemented this likelihood directly in PyMC using a pm.Potential.
One held-out PyMC example
For one untouched dataset with true
w_{\mathrm{true}}=0.70,
the hidden labels happened to contain seven M_1 fields and three M_0 fields.
If those labels had been observed directly, the oracle posterior would be
w\mid M_{1:10}
\sim
\mathrm{Beta}(8,4).
The posterior results were:
Posterior
Mean
SD
Prior Beta(1,1)
0.500
0.289
Fourier + autocorrelation ratio posterior
0.525
0.290
+ persistent homology ratio posterior
0.648
0.212
Hidden-label oracle posterior
0.667
0.131
The conventional posterior was therefore essentially unchanged from the prior.
The topology-aware posterior moved substantially toward the oracle posterior while remaining appropriately more uncertain than the oracle, since the latent labels were not directly observed.
I also independently calculated the posterior on a dense one-dimensional grid and compared it with PyMC/NUTS.
For the topology-aware model:
\text{grid mean}=0.6504,
\qquad
\text{PyMC mean}=0.6480,
and
\text{grid SD}=0.2127,
\qquad
\text{PyMC SD}=0.2123.
So the PyMC result doesn’t appear to be a sampler artifact.
Repeated parameter-recovery experiment
Finally, I froze the entire pipeline and ran a repeated recovery experiment over
w\in\{0.1,0.3,0.5,0.7,0.9\},
with 60 independent datasets per value, for 300 datasets in total.
Each dataset contained 10 noisy spatial fields.
Before running this audit I specified five success criteria.
All five passed.
Metric
Result
Conventional pooled median absolute posterior-mean error
0.218
Topology pooled median absolute posterior-mean error
0.112
Relative median-error reduction
48.6%
Topology wins paired error comparison
69.3%
Topology 90% interval coverage
95.3%
Topology / conventional median 90% interval-width ratio
0.676
Reduction in median distance to hidden-label oracle
58.4%
The recovery curves also behaved in the expected way.
The Fourier/autocorrelation posterior remained strongly shrunk toward
w=0.5,
whereas the topology-aware posterior tracked the latent mixture fraction much more closely.
Near w=0.1 and w=0.9 there is still finite-sample shrinkage toward the prior centre, which I think is appropriate given that there are only ten observed fields.
Some negative results
Not every experiment I tested worked, and I think these failures helped me clarify the use case.
My earlier attempts to improve arbitrary continuous-parameter inference using topological ABC summaries failed.
A spatial mechanism benchmark also failed because conventional geometry already separated the simulator families too easily.
A generic ABC construction based on PCA-compressed persistence images also basically returned the prior.
A learned soft-count ABC summary did move the posterior in the correct direction, but lost much of the information available to the classifier.
Those steps changed my views of the appropriate claim.
I do not think the evidence supports:
“Topology generally improves Bayesian inference.”
The narrower hypothesis I landed on would be:
Topological summaries can add useful information when a model reproduces low-order / second-order structure but misses higher-order geometric organization.
Transparency note
There was also one intermediate validation gate that narrowly failed its original stringent requirement.
The pre-specified criterion required conventional validation AUC \le 0.60.
The observed value was:
\mathrm{AUC}_{\mathrm{conventional}}=0.608.
So I do not consider that particular gate to have passed.
However, all of the incremental topology criteria passed strongly.
I therefore treated the following single-dataset likelihood-ratio inference as exploratory, froze the pipeline, and then ran the independent 300-dataset recovery audit described above.
That later audit passed all five criteria specified before evaluation.
What I am considering contributing
Given these results, I am wondering whether there is room in the PyMC ecosystem for a small, optional set of utilities around this workflow.
Such set might include:
reusable persistent-homology discrepancy / summary functions for posterior predictive checks;
representations such as persistence lifetimes, Betti curves, or persistence images;
helpers for applying topological summaries to prior/posterior predictive draws;
pm.Simulator-compatible summary functions;
examples of topology-aware likelihood-free inference;
examples of classifier-based simulation inference using topological representations.
I would prefer persistent-homology libraries such as ripser or gudhi to remain optional dependencies but am open to any recommendation.
I am also unsure where such functionality would belong.
pymc-extras seems like one plausible place for specialized functionality, but perhaps ArviZ, a small companion package, or simply additional examples would be more appropriate.
I would be happy to take on the implementation if there is interest.
Open questions on my end
The main things I would appreciate feedback on are:
Does this seem like a sufficiently useful use case to formalize beyond the existing example?
If so, where in the PyMC ecosystem would this functionality fit best?
What would you consider the smallest useful API worth prototyping first?
Would it be preferable to start with posterior-predictive diagnostics only, rather than simulator / likelihood-ratio utilities?
Would a pymc-extras prototype be a sensible first step?
I can also share the benchmark notebooks and package them into a small reproducibility repository if that would be useful.
Thanks again for the feedback on #881.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1🚀 Release pymc-extras v0.15.0019.6311-09-2026
2🚀 Release pymc-extras v0.15.1019.6316-09-2026
3🚀 Release v6.3.2018.5208-09-2026
4New contributor looking for guidance: from issue fixes to sustained PyMC contributions010.3912-09-2026
5Bambi, Count data and thresholds012.824-09-2026
6Setting and justifying priors for a discrete "what went wrong" model when I have no labeled data011.3203-09-2026
7Sampling PyMC models in JupyterLite with a WebAssembly backend for PyTensor06.3828-09-2026
8Contributions to State-Space Models & Project Ideas for PyMC-Extras014.1314-09-2026
9Introduction & GSoC 2027 Interest — PR #8434 (CAR distribution batch support)010.6516-09-2026
10Genomic epidemiology of the ongoing 2026 Bundibugyo Virus Disease outbreak in the Democratic Republic of the Congo08.0616-07-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 11.43. Источник: discourse.pymc.io.