Prism Data Lab

Every analysis here ships its dataset and the code that produced every figure. Nothing on this site is investment advice.

MethodsUpdated 1,775 words, 8 minutes

Value at Risk three ways, backtested on the S&P 500

Historical, parametric normal, and Cornish-Fisher VaR computed on ten years of FRED's S&P 500 series, then backtested day by day with the Kupiec coverage test.

Value at Risk is a quantile with a marketing department. The definition is narrow: the loss that should be exceeded only a stated fraction of the time over a stated horizon. Campbell's Federal Reserve review states it that way, and adds the sign convention that makes VaR a positive number quoted in losses. What makes VaR interesting is not the definition but the fact that three standard estimators, run on the same data on the same day, disagree by a factor of more than two. This article computes all three on ten years of a public price series, backtests them day by day, and shows where each one breaks. The data and the script are attached, and nothing here is investment advice.

The data and the loss series

The series is SP500, FRED's daily close of the S&P 500. Two things in the series notes govern what can be said with it. The source page states that "since this is a price index and not a total return index, the S&P 500 index here does not contain dividends", so every return below is a price return. And FRED carries only "10 years of daily history" for S&P and Dow Jones series, which fixes the sample window rather than any choice of ours. The file retrieved on 2026-09-08 was last updated 2026-09-04 at 19:02 CDT and ends at 7,718.60 on 2026-09-04.

That gives 2,513 daily returns from 2016-09-07 to 2026-09-04, on an index that went from 2,186.48 to 7,718.60. The loss series is simply the negative return, L = -r. Its sample statistics matter more than they usually do here: mean -0.0567 percent (negative, because the index rose), standard deviation 1.1396 percent, skewness 0.37, and excess kurtosis 15.8. Hold on to that kurtosis figure; it is what wrecks one of the three estimators.

Correction, September 18, 2026: the file attached to this article no longer contains the S&P 500 index levels. The FRED series notes carry the line "Reproduction of S&P 500 in any form is prohibited except with the prior written permission of S&P Dow Jones Indices LLC", and publishing the daily closes as a downloadable CSV was exactly that. The file now holds return_simple, the daily simple return computed from the series, which is our own output and is what every figure in this article is computed from. The index is produced by and copyright S&P Dow Jones Indices LLC. Nothing in the tables below changed: we re-ran the script against the new file and every number it prints is the number published here, and the backtest file it writes is byte for byte the one that was already shipped. To rebuild from the index itself, run python code/value-at-risk-three-ways.py --download, which pulls SP500 from FRED and writes the returns without storing the levels. FRED carries only a ten-year rolling window, so a pull today begins later than the sample reported here.

The three estimators

Historical simulation takes the empirical quantile of the loss sample. No distribution is assumed. The cost is that the estimate can only ever be as extreme as the worst loss already observed, and it moves in steps as observations drop out of a rolling window.

Parametric normal takes the sample mean and standard deviation of losses and multiplies by the inverse normal CDF at the confidence level: 1.645 at 95 percent, 2.326 at 99 percent. It is smooth and it is fast, and it assumes away exactly the property the data does not have.

Cornish-Fisher, sometimes called modified VaR, keeps the normal formula but corrects the multiplier for skewness and excess kurtosis with a fourth-order expansion:

def cornish_fisher_z(z, skew, kurt):
    return (z
            + (z * z - 1.0) * skew / 6.0
            + (z ** 3 - 3.0 * z) * kurt / 24.0
            - (2.0 * z ** 3 - 5.0 * z) * skew * skew / 36.0)

The whole script uses the standard library for the statistics, including a bisection inverse normal CDF built on math.erfc, so nothing depends on which version of a stats package is installed.

What the full sample says

Output of python code/value-at-risk-three-ways.py on 2026-09-08, one-day losses over the whole 2,513-day sample:

level historical normal Cornish-Fisher z z (CF)
95% 1.66% 1.82% 1.57% 1.645 1.428
99% 3.28% 2.59% 7.06% 2.326 6.248

Three disagreements, all of them instructive. At 95 percent the normal estimate is the highest of the three and the Cornish-Fisher estimate the lowest. At 99 percent the ordering inverts completely: the normal drops below the historical, and Cornish-Fisher produces 7.06 percent, more than double the empirical one-in-a-hundred loss. The correction has run away. An excess kurtosis of 15.8 is far outside the range in which the fourth-order expansion behaves; the adjusted multiplier of 6.248 is not a heavier-tailed quantile, it is an arithmetic artefact. That is worth stating plainly because modified VaR is often sold as the fix for fat tails, and on genuinely fat-tailed daily equity data it is capable of producing a number no one should quote.

Expected shortfall, the average loss conditional on breaching VaR, makes the tail shape visible:

level ES historical ES normal worst loss in sample
95% 2.76% 2.29% 11.98% on 2020-03-16
99% 4.75% 2.98% 11.98% on 2020-03-16

The normal ES at 99 percent is 2.98 percent. The realised average of the worst one percent of days is 4.75 percent, and the single worst day was 11.98 percent, on 2020-03-16. A normal model does not merely misprice that day, it has no vocabulary for it.

The backtest

A full-sample number is fitted on data that includes the event it is supposed to have predicted. The honest test is out of sample: estimate from a rolling window of the previous 500 trading days, forecast tomorrow, and count the days the realised loss exceeds the forecast. That yields 2,013 one-day-ahead forecasts from 2018-08-31 to 2026-09-04.

Campbell separates two properties an accurate VaR measure must have. Unconditional coverage means exceptions occur at the stated rate. Independence means one exception carries no information about the next. He is explicit that these "are separate and distinct and must both be satisfied", and that a model can pass one and fail the other. The Kupiec proportion-of-failures statistic tests only the first. It is a likelihood ratio compared with a chi-square on one degree of freedom, whose upper tail is erfc(sqrt(LR/2)) exactly, so no stats package is needed. Campbell's footnote notes that the POF test is undefined when no exceptions occur; the script handles that case with the limiting form rather than crashing.

level method expected exceptions observed rate Kupiec LR p
95% historical 100.7 105 5.22% 0.20 0.659
95% normal 100.7 102 5.07% 0.02 0.890
95% Cornish-Fisher 100.7 121 6.01% 4.08 0.043
99% historical 20.1 30 1.49% 4.25 0.039
99% normal 20.1 55 2.73% 41.44 0.000
99% Cornish-Fisher 20.1 20 0.99% 0.00 0.977

The normal model is nearly perfect at 95 percent and catastrophic at 99 percent: 55 exceptions where 20 were expected, a rate of 2.73 percent for a measure labelled one percent. That is the fat tail arriving on schedule. A distribution fitted to the middle of the data describes the middle of the data.

Cornish-Fisher inverts the pattern. It fails at 95 percent, where it is too permissive, and posts the best coverage number in the table at 99 percent: 20 exceptions against 20.1 expected, p = 0.977. It would be a mistake to read that as the model working. Its 99 percent forecasts range from 1.96 percent to 11.03 percent across the backtest, against 1.84 to 4.89 percent for historical simulation. A forecast that swings by a factor of five and lands on the right average count is not calibrated; it is unstable in a way that happens to average out. On the worst day of the backtest, 2020-03-16, the realised loss of 11.98 percent beat every forecast: historical 3.35 percent, normal 2.90 percent, Cornish-Fisher 8.20 percent.

What unconditional coverage cannot see

Exceptions by calendar year, at the 99 percent level:

year forecasts 99% exceptions (historical) 99% exceptions (normal)
2018 83 4 13
2019 252 1 5
2020 253 10 13
2021 252 0 0
2022 251 7 12
2023 250 0 0
2024 252 2 3
2025 250 6 8
2026 170 0 1

Historical simulation's 30 exceptions are not spread over eight years. Seventeen of them fall in 2020 and 2022, and two full years contain none at all. Three of the 30 landed on the day after another exception; for the normal model, 8 of 55 did. This is precisely the independence property, and the Kupiec test is blind to it by construction.

The year rows are close to the 250-day window the Basel Committee's 1996 backtesting framework uses, and Campbell reproduces the resulting multiplier bands: four or fewer exceptions at the one percent level is the green zone and five to nine the yellow, with ten or more outside both. On that scale, historical simulation lands in green in six of the nine calendar-year rows and outside it in 2020 (ten exceptions), 2022, and 2025; the normal model lands outside green in five of the nine, including 2018, 2020, and 2022. The comparison is illustrative only. A regulatory backtest applies to a bank's own trading book over a ten-day horizon, not to a one-day quantile of an equity price index, and our calendar years are not the rolling 250-day windows a supervisor would use.

Reproducibility and limits

datasets/value-at-risk-three-ways.csv holds date and return_simple, the daily simple return derived from FRED series SP500 as retrieved 2026-09-08, empty on non-trading days and on the first day of the sample. The index levels are not shipped, for the licensing reason given above; the index is copyright S&P Dow Jones Indices LLC. datasets/value-at-risk-three-ways-backtest.csv is written by the script and holds 2,013 rows: date, return_pct, loss_pct, then var95_*_pct and var99_*_pct for each of the three methods with the matching exc95_* and exc99_* exception flags (1 if the realised loss exceeded that forecast). Run python code/value-at-risk-three-ways.py --download to refresh the input; every figure above is printed by that script.

What this does not establish: nothing here forecasts a loss on any future day, and none of it is advice about holding anything. The series is a price index without dividends, so these are price losses and understate nothing on the downside but overstate nothing on the upside either. The estimation window of 500 days and the confidence levels of 95 and 99 percent are choices; a 250-day window produces different exception counts, and the script is attached so that can be checked rather than argued about. Above all, the three methods are not ranked here. Historical simulation came closest to its stated coverage at both levels, but it is bounded by its own window and it moves in steps, and one ten-year sample of one index is not a basis for a general claim.

This article is analysis and education, not investment, tax, or legal advice. Figures are cited to their source and dated; check them before relying on them.