Work MSc dissertation · 2026-09
Household responsiveness to dynamic electricity tariffs
Do households actually shift electricity consumption when prices change — and do the households least able to absorb a high price shift less than everyone else?
In one lineConsumption falls 4.16% under high prices [95% CI −4.83, −3.48]. The fairness gap I pre-registered is not there — −0.63pp, p = 0.57.
The question
Dynamic time-of-use tariffs are the main lever policy has for moving household electricity demand off the evening peak. The case for them rests on two assumptions: that households respond to price signals at all, and that the response is broad enough that the tariff does not simply transfer money from households who cannot shift to households who can.
The second assumption is the one that matters for whether these tariffs are fair, and it is the one that gets tested least. A household with electric heating, a shift worker's schedule and no dryer to defer has less room to move than a household with a dishwasher on a timer. If responsiveness tracks affluence, dynamic pricing is a regressive instrument wearing an efficiency argument.
So the fairness question is not "who uses electricity at peak" — it is "who can move it". I pre-registered that as the primary fairness hypothesis (H3): Adversity households would show a smaller absolute High-price response than Affluent households.
The data
The Low Carbon London trial, run by UK Power Networks and published through the London Datastore. Half-hourly smart-meter readings across the full trial window, 23 November 2011 to 28 February 2014, with a dynamic time-of-use tariff applied to one arm for the whole of calendar year 2013. Prices were day-ahead notified at three documented levels: High 67.20p, Normal 11.76p, Low 3.99p per kWh.
The full official release is 167,932,474 raw rows. Quality control coerced 5,560 non-numeric KWH values (0.0033%) and dropped 115,453 duplicate (household, timestamp) reads (0.069%), leaving 167,811,461 clean rows.
Three things were wrong with it, and each took real work:
The household count is contested. The London Datastore landing page says 5,567 households. The trial's own analysis reporting says 1,122 dToU. Streaming all 168 partitions gives 5,566 distinct household IDs — 4,443 standard, 1,123 dToU — and the Kaggle companion file used for the ACORN affluence mapping gives exactly the same 5,566, with a symmetric difference of zero against the official set and zero treatment-assignment mismatches across all 5,566 households. The published 5,567 is an enrolment count; one listed household contributed no half-hourly rows. I report 5,566 as the data-level count and note the reconciliation rather than quietly picking whichever number reads best. A further 5 enrolled households survive to the partition stage but contribute zero clean readings.
The tariff calendar has two sources that disagree. The official Tariffs.xlsx gives a complete 17,520-row half-hour grid for 2013 — 69 High events covering 788 half-hours, 92 Low events covering 1,660. A separate Imperial deposit (UKDS SN 7857) carries the applied prices for a different cohort and agrees on 96.3% of half-hours at zero time offset, with no candidate offset improving the match — so it is not a timezone artefact. The disagreement is concentrated almost entirely in December 2013. I adopted the official calendar as authoritative and documented why, rather than merging them.
The pre-period is not clean. Households enter the panel as their smart meters are installed, and the roll-in is steep and uneven: over 2012, standard-arm households observed per month go from 494 to 4,421, the dToU arm from 87 to 1,116. Any "2012 trend" partly reflects who was in the panel, not how people behaved.
Approach
Within-household fixed effects on the 2013 dToU panel — 18,879,437 clean half-hourly observations across 1,117 households — with half-hour-of-day, day-of-week, month and temperature controls, and cluster-robust standard errors by household.
Fixed effects rather than a pooled comparison because the heterogeneity is enormous: mean daily consumption spans roughly 0 to 101 kWh across households. Pooling would let cross-household differences in level contaminate an estimate that is supposed to be about timing. The identification comes from the same household on High half-hours versus its own Normal half-hours.
I did not run a difference-in-differences as the primary specification, and the reason is the substantive methodological content of the project. The specification, the four hypotheses, the inference strategy and a seven-rung robustness ladder were committed to a pre-analysis plan before any household-level model was estimated. Every subsequent departure is logged in a deviations file with a timestamp and a reason. The plan also contained stopping rules — conditions under which I would not proceed.
Both stopping rules fired, and the record of what happened next is the part I would most want a reader to look at:
- The falsification gate. The plan pre-committed to running the identical specification on the 4,443 never-treated standard households with placebo tariff flags, expecting approximately zero, and to stopping if the placebo exceeded 1%. It came back at +0.988% — 0.012 percentage points under the threshold, so the gate did not technically trip. But it is highly significant (p = 9.4 × 10⁻²⁹) with a confidence interval upper bound of +1.162% that does exceed the threshold. I continued, and logged it as an unresolved question rather than treating a hair's-breadth pass as a pass.
- The parallel-trends gate. A later run pre-committed to testing 2012 pre-trends before building a DiD, and to stopping if they diverged. They diverged: dToU × month dummies gave a joint Wald of 23.36 on 11 degrees of freedom (p = 0.0157), with coefficients rising monotonically through the year, and a dToU × linear trend of +2.83 × 10⁻⁴ per day (SE 7.90 × 10⁻⁵, p = 0.00034), implying +10.90% relative drift over 2012. The gate tripped and the difference-in-differences was not built. Five downstream analyses that depended on it were not run.
Inference is cluster-robust plus a design-based complement: 400 circular-shift placebo tariff calendars, which does not assume the asymptotics hold at 69 events.
The commitment, and what happened to it
What was written down before the data was looked at, and what happened to each item.
| Committed in advance | Status | Outcome |
|---|---|---|
| H1 · consumption falls under High prices | Ran | Supported. −4.16% [−4.83, −3.48] |
| H2 · consumption rises under Low prices | Ran | Supported. +3.37% [+2.87, +3.88] |
| H3 · Adversity households less responsive | Ran | Not supported. Gap −0.63pp [−2.77, +1.52], p = 0.57 |
| H4 · no positive post-event rebound | No verdict | No verdict declared — pre-trend is not flat |
| Rung 1 · temperature: linear → binned → HDD/CDH | Ran | Inside primary CI (−3.99, −4.11) |
| Rung 2 · ±1-day event placebos | Ran | ≈ 0 (−0.03, +0.04). Falsification passed |
| Rung 3 · leave-one-event-out, all 69 events | Ran | Range [−4.35, −3.79]. No single event decisive |
| Rung 4 · winsorise 99.9th, drop <90% coverage | Ran | Inside primary CI (−4.16, −4.14) |
| Rung 5 · falsification on never-treated households | Flagged | +0.99%, expected ≈ 0. Significant. Unresolved |
| Rung 6 · DST and bank-holiday exclusions | Ran | Inside primary CI (−4.31) |
| Rung 7 · Poisson pseudo-likelihood on levels | Ran | −5.35. Different functional form, same sign |
| Randomisation inference · 400 shifted calendars | Ran | p = 0.00249. 0 of 400 placebos reached the estimate |
| Gate · stop if placebo |β| exceeds 1% | Flagged | 0.988%. Did not trip by 0.012pp. Logged, not waved through |
| Gate · stop if 2012 pre-trends diverge | Gate tripped | Tripped. p = 0.0157, +10.90% relative drift |
| Difference-in-differences specification | Not run | Not built — gated by the parallel-trends failure |
| DiD event study | Not run | Not run — gated |
| H3 re-estimated under DiD | Not run | Not run — gated |
| Robustness ladder under DiD | Not run | Not run — gated |
| December exclusion under DiD | Not run | Not run — gated |
Departures from the plan, logged as they happened
- D1Ran winsorising at both the 99th and 99.9th percentilePlan specified 99.9th; an instruction specified 99th. Both reported as separate rows; neither replaces the other
- D2Added a December-2013 exclusion specificationNot in the pre-committed ladder. Added after two tariff sources were found to disagree about December. Reported as an additional row, replacing nothing
- D3Implemented the Poisson fixed-effects estimator in-houseThe standard package needs a newer Python than the pinned environment. Validated against statsmodels to 3e-15. An implementation of the committed method, not a change to it
- D4Extended the build target so a clean rebuild covers every stepA from-scratch build would have skipped the EDA and diagnostic stages
- D5Gave one figure a proper build ruleFound during a full rebuild: it had been produced by ad-hoc code and did not regenerate. Numbers unchanged
Source: ERP/Full Code/PRE_ANALYSIS_PLAN.md · DEVIATIONS.md · OPEN_QUESTIONS.md · PROGRESS.md
What I found
H1 and H2 are supported, with large and precisely estimated effects.
- High-price half-hours: −4.16% consumption (SE 0.36, 95% CI [−4.83, −3.48], p < 0.001)
- Low-price half-hours: +3.37% (SE 0.25, 95% CI [+2.87, +3.88], p < 0.001)
Both sit an order of magnitude above the pre-committed minimum detectable effect of approximately 0.48%. Randomisation inference puts both at p = 0.00249, the floor of (1+0)/(1+400) — no placebo calendar out of 400 reached the observed magnitude, and the placebo distribution for the High coefficient spans [−3.40, +3.51]%, entirely excluding the observed value.
The robustness ladder ran twelve specifications. All six genuine log-OLS variants — two temperature specifications, two winsorising thresholds, a 90%-coverage restriction, and DST/bank-holiday exclusion — plus leave-one-event-out across all 69 events stay inside the primary confidence interval. The ±1-day placebo calendars return −0.025% (p = 0.88) and +0.044% (p = 0.81): the design finds nothing when the events are moved, which is what it should do.
H3 — the fairness hypothesis, and the one I most wanted to be true — is not supported.
| ACORN group | High-price response | SE |
|---|---|---|
| Affluent | −3.71% | 0.58 |
| Comfortable | −4.66% | 0.85 |
| Adversity | −4.34% | 0.89 |
The Adversity − Affluent gap is −0.63 percentage points, 95% CI [−2.77, +1.52], p = 0.57. The point estimate runs opposite to the hypothesis — less affluent households cut slightly more, not less — and is statistically indistinguishable from zero. The null survives all eight ladder specifications, and no category among fifteen finer ACORN groups shows a gap surviving Benjamini–Hochberg correction.
Read as an equivalence statement rather than a failure to reject: the interval rules out an Adversity-less-responsive gap larger than about 1.5 percentage points. That is the honest form of the result. It is not "vulnerable households respond identically"; it is "a capacity-to-shift gap big enough to matter for tariff design is excluded, a small one is not".
A separate cohort with richer vulnerability measures points the same way. Re-running the specification on the SN 7857 deposit — a different sample, different anonymisation, its own applied price series, no crosswalk to the main sample — replicates the headline at −5.34% (SE 0.43, 95% CI [−6.14, −4.54]) and finds no significant response gap by housing tenure, heating fuel or occupancy at the 5% level. The occupancy gradient is the closest to interesting (three-or-more-person households less responsive, p = 0.055) and I have not developed it, because it did not clear the threshold I set in advance.
What this doesn't show
The headline effect is not a clean causal estimate, and I have not claimed one. The +0.988% placebo on never-treated households is the problem: High-price events were called on cold, high-demand periods, and a linear temperature control plus month and half-hour fixed effects does not fully absorb that. The event study makes the same point independently — consumption in the 24 hours before a High event is elevated by roughly 0.9–1.1%, and the pre-trend is jointly significant, so it is not flat. Both diagnostics say the same thing: some of the estimate is day composition rather than price response.
The natural fix is a difference-in-differences against the never-treated arm. I could not run it, because its identifying assumption fails its own pre-test — 2012 pre-trends diverge, and the diagnostic shows the divergence is not differential attrition (the point estimate is stable at +10.13% and +12.01% under 90% and 98% coverage restrictions; only precision falls). Staggered, arm-differential meter roll-in is a candidate explanation that I measured and did not adjudicate. So the direction of the bias is arguably known — the placebo is positive, which would make the true effect larger, not smaller — but the correction is not identified, and I would rather report a bounded estimate than a corrected one I cannot defend.
Consequently H4 has no verdict. The dynamics are consistent with conservation rather than shifting — rebound share 3.3%, cumulative effect over the event and the following 24 hours −3.56% — but a causal event study needs flat pre-trends and these are not flat. The plan asked me to report "conservation or displacement". I reported neither, and said why.
ACORN is a weak proxy for vulnerability. It is a geodemographic classification of the household's area, not a measurement of the household. Prepayment customers — the group the fairness question is really about — were excluded from the trial entirely. The SN 7857 cohort has genuine survey measures of tenure, heating fuel and occupancy and corroborates the null, which helps, but it is a different sample and cannot be merged.
December 2013 is unresolved. The two tariff calendars disagree about 87.5% of December High half-hours, against 20.3% in the rest of the year (χ² = 233.8, p = 8.8 × 10⁻⁵³). I tested and rejected my own hypothesis that constraint-management events were the mechanism — December High half-hours are the least constraint-management-tagged of any active month (10.2% against 31.4%, p = 1.3 × 10⁻⁶). The pattern looks like wholesale omission of December events from the applied-price series, but I have not established which source is correct, and it matters: excluding December moves the estimate from −4.16% to −5.18%.
The trial is thirteen years old and London-specific. Appliance stock, electric-vehicle charging and heat-pump penetration have all changed. The elasticity should not be read forward to a 2026 tariff design without argument.
What I would do next, in order: restrict the pre-period to households installed before a cut-off to remove roll-in composition and re-test parallel trends on that sub-panel; if it passes, build the DiD as a clearly-labelled secondary specification. Adjudicate the December calendar against the trial's own operational records. Then treat the occupancy gradient as a pre-registered hypothesis in a fresh sample rather than a subgroup found in this one.
The analysis is reproducible end to end. Deleting all outputs and running a single make all from raw data rebuilt 15 of 15 targets in 70 minutes with 892 values and file hashes identical to the previous run and zero substantive differences. That rebuild is also what caught a genuine gap in the dependency graph — one figure had been produced by ad-hoc code with no build rule, so it silently failed to regenerate. It is now a build target, and the fix is logged as a deviation.