MethodsUpdated 1,301 words, 6 minutes
Log returns, simple returns, and the compounding errors people make
Simple and log returns on ten years of FRED's S&P 500 series: the sum identity, variance drag, the size-of-move gap, and the mistakes each choice invites.
A return is one number, and there are two ways to write it. The simple return is the fraction the price changed by. The log return is the natural logarithm of the price ratio. For a one percent move they differ in the fifth decimal, which is why people treat them as interchangeable, and for a ten percent move they differ by half a point, which is why they should not. This article computes both on ten years of a public price series, prints the identities that hold exactly and the approximations that do not, and names the four mistakes that come from mixing them. The dataset and the script are attached.
Definitions and data
With prices P_t, the simple return is r_t = P_t / P_{t-1} - 1 and the log return is l_t = ln(P_t / P_{t-1}). They are linked exactly by l = ln(1 + r) and r = exp(l) - 1. The data is FRED's SP500 daily close, retrieved 2026-09-05, which the series notes limit to a ten-year rolling window and which excludes dividends: 2016-09-06 to 2026-09-04, 2,513 daily returns, 9.97 trading years at 252 days a year, from 2186.48 to 7718.60. Everything below is a price return.
r = pd.read_csv(CSV, parse_dates=["date"], index_col="date")["return_simple"].dropna()
l = np.log1p(r) # ln(P_t/P_{t-1}), written in terms of the simple return
Correction, September 18, 2026: the file attached to this article no longer contains the S&P 500 index levels. The FRED series notes carry the line "Reproduction of S&P 500 in any form is prohibited except with the prior written permission of S&P Dow Jones Indices LLC", and publishing the daily closes as a downloadable CSV was exactly that. The file now holds return_simple, the daily simple return computed from the series, which is our own output and is the series this article analyses. The index is produced by and copyright S&P Dow Jones Indices LLC. Nothing in the tables below changed: we re-ran the script against the new file and every number it prints is the number published here. The two index levels quoted in this article, 2186.48 and 7718.60, are the first and last closes of the sample, quoted with attribution rather than redistributed. To rebuild from the index itself, run python code/log-returns-vs-simple-returns.py --download, which pulls SP500 from FRED and writes the returns without storing the levels. FRED carries only a ten-year rolling window, so a pull today begins later than the sample reported here.
What holds exactly
Output of python code/log-returns-vs-simple-returns.py (pandas 3.0.2, numpy 2.3):
| quantity | simple returns | log returns |
|---|---|---|
| sum of daily returns | 142.53% | 126.13% |
| total return implied by the sum | 142.53% (wrong) | 253.01% (exact, as exp(sum) - 1) |
| actual total return P_T / P_0 - 1 | 253.01% | 253.01% |
| mean daily return x 252 | 14.29% | 12.65% |
| geometric annual return | 13.48% | 13.48% |
| daily standard deviation (ddof=1) | 1.140% | 1.143% |
| skewness | -0.37 | -0.68 |
The first two rows are the whole argument for log returns. The sum of 2,513 daily log returns is 126.13 percent, and exp(1.2613) - 1 is 253.01 percent, which is exactly the price ratio minus one. The sum of the simple returns is 142.53 percent, and it is not the return of anything: it is what you would have earned if you had reset to the starting capital every morning. Log returns add across time; simple returns compound, and the shortcut of adding them understates a decade's price return here by 110 percentage points.
The reverse holds across assets. A portfolio's simple return is the weighted average of its holdings' simple returns; its log return is not the weighted average of their log returns. Time-additivity and cross-sectional additivity are a trade, and no convention has both.
What is only approximate
Row four is where a quiet error lives. Multiplying the mean daily simple return by 252 gives 14.29 percent; the geometric annual return, the rate that actually turns 2186 into 7719 over 9.97 years, is 13.48 percent. Multiplying the mean daily log return by 252 gives 12.65 percent, and exponentiating that gives 13.48 percent again. The arithmetic mean of simple returns overstates the compound rate, and the gap has a name and a formula: variance drag, approximately half the variance. The script prints it two ways: the annualised difference between the mean simple and the mean log return is 1.64 percentage points, and half the annualised variance of the simple returns is also 1.64. On a series with 18 percent annual volatility the arithmetic-to-geometric gap is about 1.6 points a year, and it compounds.
The size of the gap on a single day depends on the size of the move:
| move | simple | log | gap (pp) |
|---|---|---|---|
| +1.0% | +1.00% | +1.00% | +0.00 |
| +5.0% | +5.00% | +4.88% | +0.12 |
| +10.0% | +10.00% | +9.53% | +0.47 |
| +25.0% | +25.00% | +22.31% | +2.69 |
| +50.0% | +50.00% | +40.55% | +9.45 |
| -10.0% | -10.00% | -10.54% | +0.54 |
| -25.0% | -25.00% | -28.77% | +3.77 |
| -50.0% | -50.00% | -69.31% | +19.31 |
Below a percent the two are the same number for any practical purpose. At ten percent they differ by half a point. At fifty percent the log return of a halving is -69 percent, which is the number that makes a -50 and a +50 sum to something sensible: the log returns are -0.6931 and +0.4055, they sum to -0.2877, and exp(-0.2877) is 0.75, which is exactly where a halving followed by a fifty percent gain leaves you. In simple terms the same two moves "average zero" and lose a quarter of the capital.
The largest single-day gap in the sample was 0.781 points on 2020-03-16, when the simple return was -11.98 percent and the log return -12.77 percent. The recovery needed after that day, in simple terms, was +13.62 percent, not +11.98: the asymmetry that makes drawdowns worse than they look is the same asymmetry in the table.
Why the standard deviations and skews differ
The daily standard deviations are 1.140 and 1.143 percent, a difference in the third significant figure, so any annualised volatility figure survives the choice. Skewness does not: -0.37 for simple returns and -0.68 for log. Taking the log stretches the negative tail (a -12 percent day becomes -12.8) and compresses the positive one (the +9.52 percent day on 2025-04-09 becomes +9.10), so the log distribution is more left-skewed by construction. A paper that reports return skewness without saying which return it used has reported a number that could be off by a factor of two, on this series.
The four mistakes
- Summing simple returns over time. The 142.53 percent above. Any cumulative-return column built with a running sum of percentage changes is wrong by the compounding, and the error grows with the horizon and with volatility. Use
(1 + r).cumprod() - 1or sum the log returns and exponentiate. - Averaging log returns across assets. A portfolio's return is a weighted average of simple returns. Weighting log returns gives a number that is not the portfolio's log return, and the error again scales with the dispersion of the holdings.
- Annualising the arithmetic mean and calling it the compound rate. The 14.29 percent versus 13.48. The arithmetic mean is the right input to a Sharpe ratio's numerator under the usual conventions, and the wrong answer to "what rate did this grow at".
- Comparing statistics computed on different conventions. Skewness, kurtosis, maximum-loss thresholds, and anything involving a large move differ materially between the two. State the convention in the header of the table, every time.
What is not settled
Neither convention is "correct". Simple returns are what a brokerage statement shows and what portfolio arithmetic needs. Log returns are what continuous-time models assume, what makes multi-period sums exact, and what makes returns closer to symmetric in distribution (though not, as the -0.68 skew shows, actually symmetric). The choice is a modelling decision with consequences that are small for daily moves and large for anything cumulative or extreme, and the script prints every number here so the consequences can be checked on any other series by pointing it at a different CSV.