Toby Lowe Analysis you can check · Manchester

Work MSc coursework · 2026-05

UK banking and payments to 2034

What can a twenty-five-year record of UK banking and payments behaviour honestly support as a projection to 2034 — and where does the projection stop being defensible?

In one lineBranch network down over 70% since 2000, r = −0.98 against remote banking. The 2034 figures are scenarios, and the assumptions dominate them.

Attribution
Group project of five (P10, DATA70202, in partnership with FIS Global). My contribution per the team-roles table — data collection for smartphone ownership and contactless payment limits; modelling; report writing and final integration.
Methods
Bass diffusion · Fisher-Pry substitution · OLS · scenario forecasting · sensitivity analysis
Tools
python · pandas · statsmodels

The question

FIS Global asked for an assessment of long-term structural shifts in UK banking and payments that could affect payment processing, merchant services, banking platforms and fraud prevention.

The interesting constraint is not the domain, it is the sample. Twenty-five years of annual public data is 25 observations. For the payment-method series it is worse — reliable coverage starts in 2014, and removing 2020 as an outlier leaves n = 10. Any model with a meaningful parameter count will fit that beautifully and forecast nothing.

So the real question underneath the brief was a methodological one: what is the most that can be claimed from ten to twenty-five annual observations, and how do you present a 2034 number without implying a precision you do not have?

The data

Public secondary sources only — Bank of England, UK Finance, ONS, Pay.UK, Ofcom and the LINK Scheme. No FIS internal data and no individual-level customer records were accessed at any point, which was a scoping decision made at the outset rather than a limitation discovered later.

The team ran a source audit before modelling, classifying every variable in the brief as used directly, used contextually, or excluded, with a stated reason. That audit is where the honest weaknesses are recorded:

  • Coverage is not uniform. Banking infrastructure and interest rates cover the full period. Payment-method indicators only run from 2014.
  • Some requested variables have no usable series at all. Mobile wallets, wearables and neobank primary-current-account share lack a time series entirely — so anything said about them in 2034 is a scenario constructed from adoption theory, not an extrapolation of observed data.
  • Regional analysis is impossible longitudinally. No comparable regional series exists across the study window; cross-sectional FCA data was substituted, which answers a different question.
  • Average account-opening time was snapshot-only and too sparse to forecast, so it was used descriptively and excluded from modelling.
  • Smartphone ownership (27% → 95%, 2011–2024) was treated as an exogenous enabling driver rather than merged as a time series — one of the two datasets I collected, and treating it as a driver rather than a regressor is the reason it does not contaminate the sample size.

2020 is a genuine structural break rather than a nuisance outlier, and the analysis treats it as one of three identified inflection points alongside post-2008 branch rationalisation and mid-2010s contactless adoption.

Approach

Parameter-light models, chosen because of the sample size rather than in spite of it. With n between 10 and 25, the modelling choice is essentially forced: three-parameter diffusion models with theoretically-interpretable parameters, not flexible learners.

  • Bass diffusion for mobile-wallet penetration — an adoption process with innovator and imitator terms, applied where the quantity is genuinely a new-product adoption curve.
  • Fisher–Pry substitution for neobank share of primary current accounts — a logistic displacement model, applied where the quantity is one technology taking share from a functional competitor. Using the right S-curve for the right structural story rather than fitting a generic logistic to both.
  • Floored exponential decline for cash transaction volumes, with the floor motivated by evidence rather than fitted — the Access to Cash Review's finding that 17% of UK adults remain significantly dependent on cash implies a structural floor that an unconstrained exponential would forecast straight through.
  • OLS for the macroeconomic relationship between Bank Rate and deposit balances.

Everything is reported as Low / Base / High scenarios, with an assumptions register and a sensitivity analysis, rather than as point forecasts.

The methodological decision I would most want to be asked about at interview happened in the data-architecture sprint, and it was resolved out of a genuine disagreement in the team. One proposal was a single wide master table keyed on year — clean, and it would have simplified the correlation analysis, the regression and all downstream feature engineering. The competing proposal was three thematic dataframes merged ad hoc per analysis.

The wide table was rejected because a full merge would have imposed the shortest series' range on every variable, collapsing the effective sample from n = 25 to n = 10 on every cross-dataset operation. The convenient architecture would have quietly destroyed more than half the data. The ad-hoc approach cost real engineering time and preserved the OLS sample at n = 25 (R² = 0.627).

That is the general principle the project ran on: where data-engineering convenience and statistical defensibility pull in opposite directions, the sample wins.

What I found

The historical record supports three findings firmly, because they rest on long series rather than short ones:

  • Bank branch volume is down over 70% since 2000, correlating at r = −0.98 with the rise in remote banking and faster-payments volumes.
  • The rate–deposit relationship persists, at r = 0.792 between Bank Rate and personal deposit balances.
  • Fraud has migrated rather than reduced. Counterfeit-card losses fell 91% between 2012 and 2024 as physical-card controls matured, with card-not-present becoming the primary vector.

The 2034 figures are scenario outputs, not predictions, and the report says so explicitly. Under the Base scenario:

Projection (2034, Base scenario)Method
~85% mobile-wallet penetrationBass diffusion
~49% neobank share of primary current accountsFisher–Pry substitution
~2.6bn cash transactions (structural floor)Floored exponential decline

Each of these rests on a series that either starts in 2014 or does not exist as a time series at all. The sensitivity analysis confirms that the structural ceiling and floor assumptions — not the fitted parameters — are the dominant source of forecast uncertainty. In other words, the numbers are mostly determined by the judgement calls declared in the assumptions register, which is precisely why the register exists and why these should be read as bounded scenarios rather than estimates.

The strategic recommendations that follow — prioritising real-time payment orchestration, Banking-as-a-Service infrastructure, and AI-enabled fraud detection for card-not-present vectors — follow from the direction of the findings, which is well supported, rather than from the magnitude of the 2034 figures, which is not.

What this doesn't show

Nothing here is causal. Every cross-dataset result is an association on annual aggregates. Branch decline, faster-payments growth and cash decline all have external drivers — merchant acceptance, regulation, technology cost — that are not in any model. The r = −0.98 between branch closures and remote banking is two trends over the same period, and a correlation that high on 25 annual observations of two trending series should be read as description, not mechanism.

The 2034 figures should not be quoted as forecasts. They are scenario outputs whose values are driven mainly by declared floor and ceiling assumptions. Quoting "85% mobile-wallet penetration by 2034" without "Base scenario, Bass diffusion, no underlying time series for this variable" attached would misrepresent the work. This is the claim on the whole site I would be most careful about restating.

The sample is too small for the usual defences. Parameter-light modelling protects against overfitting; it does not create information. There are no held-out years, no backtest, and no out-of-sample validation, because there is not enough data to hold any out.

No disaggregation is possible. Public annual aggregates cannot be broken down by region, age, income, merchant category or channel, so nothing here says anything about who is being left behind by cash decline — which is the question with the most policy weight and the one the data cannot reach.

Attribution. This was a five-person project and the findings are the team's. My own contribution was the smartphone-ownership and contactless-limit datasets, modelling work, and report writing and final integration. The data-architecture decision described above was a team decision reached through disagreement, not mine alone.

What the team identified as next steps: treat the 2026 contactless cap removal as a natural experiment on contactless usage — a genuine identification opportunity rather than another trend fit; obtain segment-level transaction data, suitably anonymised, to convert national trends into something disaggregated; and once longer post-2024 series exist, benchmark these interpretable baselines against Bayesian structural time-series or state-space alternatives, which is the comparison that would actually test whether parameter-light was the right call.