Work MSc coursework · 2026-05
Persistence norms and conditional volatility in European gas markets
Does the topological shape of gas-price returns carry regime information that GARCH-family volatility models cannot see?
In one line26 sustained regime signals at the crisis peak against GARCH's 0 — and 0 against 39–46 through the recovery. Complementary, not superior.
The question
Regime detection in energy markets is dominated by GARCH and its extensions. They model the magnitude of return variation, and they are good at it. But magnitude is not the only thing that changes when a market enters a crisis — the geometry of the return dynamics changes too, and a variance model has no vocabulary for that.
Persistence norms, computed from Takens embeddings of a return series, do have that vocabulary. They summarise how much cyclic structure exists in the reconstructed state space. The literature establishing them as market-stress indicators is concentrated in equities and cryptocurrency, and rests on qualitative inference over a small number of crashes. Energy-market regime detection, by contrast, leans heavily on GARCH and regime-switching variance models, which leaves the value of topological signals in this setting an open question — and that is the gap this work takes on.
The question is deliberately not "is TDA better than GARCH". It is the sharper, more falsifiable one: is there regime information in the topological signal that the variance signal misses, and is there information in the variance signal that the topological signal misses? A method that only ever agrees with the incumbent is not worth the compute.
The data
TTF Dutch natural gas futures daily close prices, January 2019 to April 2026, in EUR/MWh, from Yahoo Finance — N = 1,843 observations, analysed as log returns.
The sample is deliberately built around two supply-side shocks: the Russia/Ukraine crisis (September 2021 – January 2023, with TTF peaking near EUR 340/MWh in August 2022) and the Iran/Hormuz disruption (February 2026 onwards, following Qatari LNG flow disruption, with TTF reaching roughly 200% of its post-crisis baseline).
Two things about the data constrain what can be concluded. The return distribution is heavily fat-tailed — standard deviation 0.0542 against a range of −0.3524 to +0.4128 — which is part of the motivation for a non-parametric method but also means any Gaussian-flavoured diagnostic should be treated carefully. And the COVID period sits inside the sample as a demand-side shock that is not one of the crises under study; both detectors fire during it (22 days for the topological signal, 16 for GARCH), which is noise for this question rather than a finding.
The second crisis window is only 62 observations. It is a forward-looking check, not evidence.
Approach
Takens embedding parameters were selected by standard, non-discretionary rules rather than tuned: time delay τ = 2 as the first local minimum of average mutual information (Fraser & Swinney), and embedding dimension m = 4 as the smallest dimension at which the false-nearest-neighbour fraction falls below 1% (Kennel et al.). The primary signal is the dimension-one L¹ persistence norm — the total lifetime of one-dimensional homological features, i.e. how much loop structure exists in the embedded point cloud.
The comparator ladder is deliberately unfavourable to the new method: GARCH(1,1), GJR-GARCH(1,1) to allow for leverage asymmetry, and a two-regime Markov-switching variance model. If a topological signal only beats a plain GARCH(1,1), that is a weak claim; it needs to beat the model class.
The detection rule is the same for every signal, which is the design decision that makes the comparison fair. Each signal is converted at each date to its empirical CDF rank against the preceding 252 trading days — a probability integral transform — and a regime change is flagged when that exceeds z > 2 for three consecutive days. Same threshold, same window, same persistence requirement, so any difference in detection is a property of the signal and not of the rule applied to it.
Two protocols were run: level-based detection on the raw signal, and innovation-based detection on AR(1) residuals for the persistence norm and standardised residuals for the GARCH family. The second protocol is the one that tests whether an apparent advantage survives letting each signal adapt to its own dynamics.
What I found
The two signals carry complementary regime information, and the complementarity is phase-specific.
Sustained exceedance days, level-based protocol:
| Period | L¹ persistence norm | GARCH(1,1) | MS-GARCH |
|---|---|---|---|
| Crisis 1 peak (Sep–Dec 2022) | 26 | 0 | 0 |
| Recovery (2023–2025) | 0 | 39 | 46 |
At the Crisis 1 peak the persistence norm produced 26 sustained exceedance days and GARCH(1,1) produced none. This holds against the stronger comparators: GJR-GARCH also yields zero plateau exceedances (γ̂ = −0.017, p = 0.57), and the Markov-switching model yields zero despite correctly identifying the high-variance regime on 87% of plateau days (mean P(high) = 0.982).
That last detail is the actual finding, and it is a mechanism rather than a scoreboard. The variance model was not wrong — it saw the high-variance regime. It produced no trigger because volatility was not unusual relative to its own rolling history once the elevated level had persisted. The suggestion is that the Crisis 1 plateau was characterised by a change in the structure of return dynamics rather than by exceptionally large price movements.
In the recovery period the result inverts: the persistence norm returns to baseline with zero exceedance days while GARCH and MS-GARCH register 39 and 46. This replicates, in energy commodities, the finding from the equities literature that norms retreat to baseline faster than variance-based measures.
The asymmetry is a property of level-based detection specifically. Under the innovation-based protocol the Crisis 1 gap narrows from 26–0 to 7–1. That is a substantial weakening, and it localises the finding: once each signal adapts to its own elevated dynamics, topology and variance largely coincide. The honest reading is that the topological signal adds most at the regime level, not to short-horizon forecast errors.
Two robustness checks support the specification without rescuing it beyond its strength. Embedding dimension sensitivity behaves as the selection rule implies — m = 2 flattens the signal entirely (0 exceedances at the Crisis 1 peak), m = 3 and m = 4 fire selectively (22 and 26), and m = 5 fires during otherwise calm periods (20 post-crisis exceedances against 0 at m = 4), consistent with ambient-dimension noise. A placebo on Henry Hub gas over a calm 2010–2017 window has both detectors firing at approximately the nominal 5% rate, so the pipeline does not manufacture exceedances from quiet data.
On the forward-looking crisis, GARCH triggered first (3 March 2026) and the persistence norm second (30 March 2026). That is one event and I draw nothing from it.
The conclusion the evidence supports is the modest one the report states: persistence norms should be used alongside conditional variance measures, not instead of them.
What this doesn't show
Two crises is not a sample. Every number here is a description of two episodes in one commodity. The Crisis 2 window is 62 observations. Nothing in this design supports a claim about how the detector would behave on the next shock, and the fact that the two crises produced opposite orderings on first trigger should be read as exactly that instability.
Detection is not prediction. The protocol flags regime changes contemporaneously against a trailing 252-day baseline. It is not a forecast, was not evaluated out-of-sample in any walk-forward sense, and no trading or hedging value is claimed or tested.
The detection threshold is a free parameter and I only ran one. z > 2 with a three-day persistence requirement is conventional, and it is applied identically to every signal so the comparison is internally fair — but the magnitudes (26 versus 0) would move under a different threshold, and I have not mapped how much. A sensitivity sweep over the threshold and the persistence requirement is the first thing this needs.
The COVID window contaminates the calm periods. Both detectors fire during it. It is a demand-side shock rather than one of the supply-side crises under study, and it sits inside what would otherwise be baseline. It is acknowledged rather than handled.
The mechanism is inferred, not demonstrated. "The market changed shape rather than amplitude" is a reading consistent with the Markov-switching result, not something the analysis establishes independently. Testing it properly means relating the persistence norm to observable microstructure — contract roll behaviour, curve shape, storage levels — rather than to another return-based statistic.
Yahoo Finance is a convenience source for a front-month futures series, and continuous-contract construction can introduce roll artefacts at the exact moments a topological method is most sensitive to. An exchange-sourced series with an explicit roll convention would be the correct data.
What I would do next: sweep the threshold and persistence rule; extend to a second commodity with independent crisis timing to get event count above two; and test the persistence norm against a structural-break test rather than a volatility model, since the claim being made is really about structural change, not variance.