Here is a bet. You put up 100 dollars. A fair coin decides. Heads, your money grows by 50 percent. Tails, it shrinks by 40 percent. Expected value per round: plus 5 percent. Take the bet, and take it again, and take it forever. What happens?
The intuitive answer, the one every finance textbook and every MBA calculator will hand you, is that you get rich. Compound plus 5 percent for long enough and you compound your way to a fortune.
The correct answer is that you go broke, almost surely, and the math is not close.
This is not a paradox. It is a theorem, and it has been proved and reproved for three centuries under different names. It is why Daniel Bernoulli wrote the first paper in mathematical finance in 1738. It is why the Kelly criterion exists. It is the single most robust reason the average investor systematically underperforms the average return. And it comes down to one fact that almost nobody teaches: there are two different averages in probability, they answer two different questions, and mainstream finance overwhelmingly uses the wrong one.
The technical name for the mistake is a failure of ergodicity. The practical name is the reason your account keeps losing money on bets you can prove are winners.
The bet you cannot win by winning it
Take the coin flip above and let it run.
If you multiply out the two possible outcomes you get a growth factor per round. Heads gives 1.5, tails gives 0.6. If you play many rounds and land roughly half of each, your wealth ends at 1.5^n × 0.6^n = 0.9^n dollars per dollar started. That is not growth. That is a 10 percent loss per round pair, compounded.
The geometric mean of 1.5 and 0.6 is the square root of 0.9, which is about 0.9487. So the return you actually experience, on the actual path you actually walk, is negative 5.13 percent per round.
Simulate ten thousand gamblers running this bet a hundred times each and the divergence is not subtle.
bash
import numpy as np
rng = np.random.default_rng(0)
paths = np.ones((10_000, 101))
for t in range(1, 101):
flip = rng.random(10_000) < 0.5
paths[:, t] = paths[:, t-1] * np.where(flip, 1.5, 0.6)
print(f"Ensemble mean at t=100: {paths[:, 100].mean():>12.2f}")
print(f"Median gambler at t=100: {np.median(paths[:, 100]):>12.5f}")
print(f"Share of gamblers broke: {(paths[:, 100] < 0.01).mean():>12.1%}")
Ensemble mean at t=100: 131.24
Median gambler at t=100: 0.00584
Share of gamblers broke: 88.7%Both numbers are correct. Both are computed from exactly the same distribution. The 131 is the average across the ensemble. The 0.006 is the median individual outcome, which is what you would see if you were one of the ten thousand gamblers. Ten thousand ran the bet, roughly nine thousand lost almost everything, and a handful of astronomical winners in the right tail dragged the mean above the starting stake.
If you take the bet expecting the plus 5 percent number to accrue to you, you are betting on being one of the winners in the tail. If you take it expecting to compound your money, you are betting against the math.
The two averages
The confusion has a name. A process is called ergodic if the average across many copies at one moment equals the average of one copy over a long time. For a fair, additive process this is true. Roll a die a million times and roll a million dice once, you get the same distribution.
For a multiplicative process, which is what compounding money is, it is false. The ensemble average and the time average diverge, and the gap widens with volatility.
Ole Peters, at the London Mathematical Laboratory, has spent a decade dragging this point back into finance in the language of physics. His observation is not new. Bernoulli anticipated it in 1738, Kelly formalized a version of it in 1956, and the entire practice of maximizing log utility is a workaround for the same underlying issue. But the ergodicity framing makes it precise: a great deal of modern finance quietly assumes ergodicity for processes that are demonstrably non-ergodic, and treating them as if they were is the source of a specific, calculable, recurring loss.
The mathematical form is clean. Under geometric Brownian motion, the standard model of asset prices,
dS/S = μ dt + σ dW
the ensemble average grows at rate μ. The time average grows at rate
g = μ − σ² / 2
The gap is exactly σ²/2, a subtraction called the Itô correction, or in finance the volatility drag. It is not an approximation. It is not a behavioral finding. It is what the stochastic calculus of a multiplicative random process says.
The subtraction that decides everything
Look at the correction across realistic parameters and it stops being a curiosity.
┌──────────────────┬──────────────────┬────────────────────────┐
│ Expected μ │ Volatility σ │ Time-average g │
├──────────────────┼──────────────────┼────────────────────────┤
│ 8% │ 20% │ +6.0% │
│ 8% │ 30% │ +3.5% │
│ 8% │ 40% │ 0.0% │
│ 8% │ 50% │ −4.5% │
│ 15% │ 80% │ −17.0% │
└──────────────────┴──────────────────┴────────────────────────┘The first row is roughly the S&P 500. The second is an aggressive equity or emerging-markets index. The third is a mid-cap growth stock or a modestly leveraged position, where the advertised expected return of 8 percent produces a realized growth rate of zero. The fourth is small caps and much of the crypto market, where a positive expected return still hands you a compounded loss. The fifth row is where meme coins and single-stock moonshots sit, and the realized growth is deeply negative even though the pitch deck screams positive alpha.
This is why the average investor in a fund earns systematically less than the fund's advertised return. It is why leveraged ETFs marketed as 2x the S&P or 3x QQQ almost never deliver 2x or 3x their underlying over any long horizon, and in flat markets they lose money outright. The industry writes the fund brochure in ensemble averages. The investor lives in time averages. The gap is the σ²/2 term, and the gap is where the money goes.
What Kelly was really solving
The Kelly criterion, published by John Kelly in the Bell System Technical Journal in 1956, is usually taught as a bet-sizing rule. What it actually is, mathematically, is the position size that maximizes the time-average growth rate of wealth given the returns you have and the odds you face.
For a position sized at fraction f of wealth in an asset with expected return μ and volatility σ, the resulting time-average growth rate is
g(f) = f·μ − f²·σ² / 2
Maximize with respect to f and you get the continuous-time Kelly optimum
f* = μ / σ² g* = μ² / (2·σ²)
Written this way, the equivalence to ergodicity is exact. Kelly asks: how much of my wealth should I risk on each opportunity so that my long-run growth rate along my one actual path is as fast as possible? The answer is the fraction that maximizes the expected value of log(1 + r), which is the ergodic transformation of the naive expected value. Log utility is not a psychological assumption about human preferences. It is the correct utility function for anyone whose wealth compounds multiplicatively and who only gets to live in one universe.
Every serious practitioner who has grown a fortune from a small edge, Thorp, Simons and the Renaissance team, Bill Gross in his card-counting days, has been solving the same ergodicity problem, whether or not they used the word. They size for time-average growth. They discount naively expected return by the volatility drag. They accept a lower headline number in exchange for growth that actually accrues to their account rather than to an imaginary ensemble.
What the correction actually looks like at the desk
The most instructive place to see the math bite is leverage. Suppose you have identified an asset with μ = 8 percent and σ = 20 percent. Kelly says the optimal leverage is f* = 0.08 / 0.04 = 2.0. Two times leverage maximizes your long-run growth rate. Push past that number and every additional unit of leverage lowers your realized growth even as it raises your ensemble mean return.
Watch what leverage does at each level, all for the same underlying asset.
┌───────────────┬──────────────┬──────────────┬─────────────────┐
│ Leverage f │ f·μ │ f²·σ² / 2 │ Growth g(f) │
├───────────────┼──────────────┼──────────────┼─────────────────┤
│ 1.0 │ 8.0% │ 2.0% │ +6.0% │
│ 2.0 │ 16.0% │ 8.0% │ +8.0% ★ │
│ 3.0 │ 24.0% │ 18.0% │ +6.0% │
│ 4.0 │ 32.0% │ 32.0% │ 0.0% │
│ 5.0 │ 40.0% │ 50.0% │ −10.0% │
└───────────────┴──────────────┴──────────────┴─────────────────┘The 2x Kelly-optimal position grows at 8 percent per year, forever. The 3x position grows at 6 percent, the same rate as the unlevered position, because the extra leverage has added exactly as much volatility drag as it has added expected return. The 4x position stops growing entirely. The 5x position destroys the account at 10 percent per year even though its ensemble expected return is 40 percent. The parabola is not a curiosity of the equation. It is the ledger of the account.
This is the mathematical fence behind every fund that survives a decade and behind every over-leveraged retail account that does not. Discovery of the fence is not a matter of intelligence. It is a matter of writing g(f) down and looking at where its derivative crosses zero.
Beyond leverage, the operational corrections stack:
Convert every reported return to a geometric growth rate before comparison. For a fund reporting 12 percent expected return at 25 percent volatility, the time-average is 12 minus 3.125, or 8.875. That is what accrues to a buy-and-hold investor in that fund.
Optimize allocations for time-average growth, not expected return. Two portfolios with identical expected returns can produce materially different geometric growth rates depending on their volatilities and correlations. Preferring the lower-variance configuration at the same expected return is not conservatism. It is arithmetic.
Discount any strategy whose track record is short and whose returns are volatile. Short samples in high-variance environments are dominated by ensemble-average estimation error. Realized returns are the naive expected value plus noise, and for a strategy with true edge close to zero, most of what you are extrapolating is noise. Lo (2002) and Bailey and López de Prado (2014) formalize this on the Sharpe side. It composes with the ergodicity correction rather than competing with it.
What the industry knows and does not say
None of this is hidden. Peters has been publishing on ergodicity economics since 2011. Kelly's paper has been continuously in print since 1956. Bernoulli's paper on the St. Petersburg paradox, which is essentially the same discovery in older language, has been available for close to three centuries.
What the industry does with the knowledge splits cleanly. The serious quantitative shops price ergodicity carefully. They compute geometric growth rates rather than expected returns. They size positions with the volatility drag included. They think in log utility whether or not they use the word. They have to, because they are compounding real money over long horizons and they have discovered by natural selection that the naive expected-value calculation eventually blows up.
The retail side of the industry does the opposite. Every fund advertisement, every robo-advisor projection, every retirement calculator you have seen uses arithmetic expected returns. The number is larger, the story is simpler, and the client is happier at the point of sale. That the number is systematically wrong for the client's actual outcome is not disclosed, because there is no rule requiring it and no incentive to volunteer it. The client will believe they underperformed the market when in fact they earned exactly the time-average growth rate the math predicted from the day they bought in.
The one thing to remember
There are two averages in probability. One counts what happens across all the versions of you that could have existed. The other counts what happens to the one version that does. For any additive process, they are the same number. For any multiplicative process, they are different numbers, and the difference grows with the variance.
Your wealth is a multiplicative process. Your leveraged wealth is a more severely multiplicative process. Anything that compounds in your life, including reputation, capital, skill, and health, obeys the same math. In each case the number the world quotes at you is not the number that runs your account.
The correction is small in symbols and enormous in dollars. Subtract σ²/2 before you get excited about a return. Understand that the fund brochure is describing a different investor than you, one who lives in every parallel universe at once. You live in only one, and you are entitled to a smaller and more honest number.
The industry has been quoting the wrong average for a long time. The math for the right one has been in plain sight for nearly three hundred years. Nothing about the situation is hidden. The whole problem is that the wrong number is larger, easier to explain, and much more comfortable to believe. It just happens to describe an outcome that will not arrive at your account.