Knowledge · Research · Alpha decay: what dies out-of-sample, what survives, and why

Alpha decay: what dies out-of-sample, what survives, and why

Strategy analysis 2026-07-25 13 sources
The evidence-based answer to the deepest doubt of the whole project: are there any setups that actually work, has all the alpha in liquid assets already been arbitraged away, and do backtests even make sense? A deep-research pass over 23 peer-reviewed findings (JF/RFS/JFE/NBER) draws one sharp line: data-mined patterns decay reliably after publication, structural risk premia persist because they are paid compensation for real crash risk. This is the positive companion to 'What provably does NOT work' and it independently reproduces the exact distinction Botty's own two years of backtests forced on us.
  • TWO KINDS OF EDGE. The literature splits cleanly. Statistical anomalies (data-mined against the cross-section of returns, no economic payer) decay. Structural risk premia (carry, variance risk premium, trend) persist because they compensate you for bearing real, correlated, crash-timed risk. Botty's survivors are all of the second kind; every dead TA/pattern signal was of the first.
  • ANOMALIES DECAY, QUANTIFIED. McLean & Pontiff (2016, Journal of Finance), 97 predictors: portfolio returns fall 26% out-of-sample and 58% post-publication. Decay is gradual (~25% at 3yr, ~40% at 5yr, ~50% by 10yr). The extra ~32 points between OOS and post-publication is arbitrage capital removing the mispricing once it is public.
  • MOST PUBLISHED FACTORS FAIL REPLICATION. Hou-Xue-Zhang (2020, RFS): of 452 anomalies, 65% fail t>1.96 and 82% fail t>2.78 under robust construction; survivors are much smaller than originally reported. Harvey-Liu-Zhu (2016, RFS): a new factor should clear t>3.0 given the scale of data-mining. This is exactly the Deflated-Sharpe / t>3 bar Botty's indicator_lab already enforces.
  • THE COUNTER-CAMP (do not ignore). Jensen-Kelly-Pedersen (2023, JF) find 55-61% of asset-pricing factors DO replicate and are strengthened, not weakened, by the number of correlated factors. Chen-Zimmermann show only ~half of the observed decay is publication bias (bias-corrected drop only ~12%). The replication rate is a genuine, methodology-dependent dispute. But note: this rescues the structural survivors, not dead price patterns.
  • CAPACITY IS THE RETAIL EDGE (the hard proof). Jacobs & Müller (2020, JFE): across 241 anomalies in 39 markets, the US is the ONLY country with a reliable post-publication decline. Edges die where cheap arbitrage capital removes them, not everywhere. Corollary: small/illiquid markets keep edges that are uneconomic to harvest at institutional scale, not hidden. A solo trader's advantage is capacity, not information.
  • CARRY IS A DURABLE, DIVERSIFIABLE RISK PREMIUM. Koijen-Moskowitz-Pedersen-Vrugt (2018, JFE): carry predicts returns across equities, bonds, FX, commodities, credit; within-class Sharpe ~0.7, diversified across classes ~1.1 with near-zero cross-class correlation. It persists because the three biggest carry drawdowns (1972-75, 1980-82, 2008-09) coincide with global recessions. Crypto funding carry is one sleeve of this premium, not a standalone free lunch.
  • VRP / SHORT-VOL IS PAID FOR TAIL RISK. Heston-Todorov: 19 of 20 futures across asset classes have a negative variance risk premium. Bollerslev-Todorov: roughly half of the S&P 500 VRP is direct compensation for jump/disaster risk (left tail 59.8% vs right tail 10.0% of the premium). Profitable most years, structurally short the tail. This is exactly Botty's VRP-harvest profile.
  • TREND: BREADTH, NOT DEPTH (the strongest strategic lever). AQR 'Trends Everywhere': a diversified multi-market trend strategy has gross Sharpe ~1.60 versus ~0.35 for the median single instrument. Hurst-Ooi-Pedersen (century evidence): net-of-fees Sharpe ~1.0, positive in every decade 1903-2012. Botty's single-asset BTC-Donchian IS the 0.35 version. The signal is fine; the noise was never diversified away.
  • POST-2009 HONESTY. Fast trend-following collapsed after ~2009 (EWM-5-20 Sharpe 0.84 to 0.12), slow trend held up far better (EWM-50-200 0.70 to 0.40). Treat the century Sharpe as a historical upper bound; the deployable version is slow, diversified, and cost-aware.
  • FRICTIONS DESTROY SMALL-CAP CRYPTO SIGNALS. The documented crypto low-idiosyncratic-risk anomaly is statistically present (-1.11% weekly) but non-exploitable after transaction costs and margin-call-driven drawdowns (Pacific-Basin Finance J. 2023). Directly relevant to expanding funding carry into smaller perps: the fatter nominal edge sits exactly where slippage, delisting and liquidation risk are largest.
  • DO BACKTESTS STILL MAKE SENSE? Yes, but as a FALSIFIER, not a DISCOVERER. With enough trials, noise guarantees a beautiful backtest (Botty's own megasweep PBO ~0.76 is the mathematical proof). The fix is order-of-operations: ask the economic question first (who pays me, and for what risk?), then backtest to verify and size. That order is exactly what separates Botty's survivors from its dead signals.
P1 Treat multi-market trend (breadth, not depth) as the evidence-strongest new direction, above any further single-asset BTC entry search.
The 1.60-vs-0.35 diversified-vs-single Sharpe gap (AQR) and century-long positive-every-decade record (Hurst-Ooi-Pedersen) are the best-documented systematic edge there is. Botty's BTC-Donchian is the low-Sharpe single-asset version by construction; the payoff is in breadth. Constraint: slow signals only (fast trend collapsed post-2009), and it needs multi-market futures/perp data + broker infra.
P2 Before expanding funding carry into smaller/illiquid perps, bench the net-of-cost, net-of-delisting yield in-house and gate hard on it.
Capacity theory says the fatter funding on small coins is real and retail-harvestable. But the crypto friction evidence (low-idio anomaly present at -1.11%/wk yet non-exploitable after costs + margin-call drawdowns) warns those edges live exactly where slippage, delisting and liquidation risk bite hardest. The nominal edge is not the deployable edge.
P3 Keep the economic-rationale-first discipline as the gate before any backtest: name the risk premium and the payer, or do not build.
PBO ~0.76 proves the megasweep can manufacture beautiful noise. A backtest is a falsifier, not a discoverer. Every Botty survivor answers 'who pays me and for what risk'; every dead signal could not. This ordering is the cheapest overfit defense there is.

What this is about

After two years and tens of thousands of backtests, essentially every TA / pattern-based BTC-timing signal Botty tested died out-of-sample (walk-forward + PBO ~0.76), while a handful of structural bets (funding carry, variance risk premium, cross-sectional momentum, trend) survived. That raised the deepest possible doubt: are there any setups that actually work? Has all the alpha in the big liquid assets already been eaten? Do backtests even make sense?

This page is the answer, sourced from peer-reviewed finance (Journal of Finance, RFS, JFE, NBER) plus credible practitioner research (AQR). It is the positive companion to What provably does NOT work - and the striking result is that the academic literature independently draws the exact same line Botty's own data drew.

The one distinction that explains everything

There are two kinds of "edge," and they behave completely differently:

1. Statistical anomalies - a pattern data-mined against the cross-section of returns, with no story for who is paying you. These decay reliably once published, because there is no economic reason for them to persist and arbitrage capital removes them.

2. Structural risk premia - you are paid to bear a real, correlated, crash-timed risk that other people want to offload. These persist out-of-sample, precisely because the premium is compensation, not a free lunch. It has to hurt sometimes, or it would not exist.

Every dead Botty signal (divergences, Wyckoff/SMC/ICT, OPEX/max-pain, sentiment-contrarian, volume profile, cross-sectional reversal, single-asset vol-targeting, the whole megasweep of TA entries) was of the first kind. Every survivor (funding carry, VRP, cross-sectional momentum, trend) is of the second. This was not luck; it is the predicted outcome.

Anomalies decay - the numbers

  • McLean & Pontiff (2016, JF): 97 predictors decline 26% out-of-sample and 58% post-publication. Gradual: ~25% at 3yr, ~40% at 5yr, ~50% by 10yr. The gap between the two figures is arbitrage capital trading on the now-public signal.
  • Hou-Xue-Zhang (2020, RFS): of 452 anomalies, 65% fail t>1.96, 82% fail t>2.78 under robust (value-weighted, microcap-controlled) construction; survivors are far smaller than first reported.
  • Harvey-Liu-Zhu (2016, RFS): given the scale of data-mining, a new factor should clear t>3.0. This is exactly the Deflated-Sharpe / high-t bar Botty's indicator_lab already runs.

The honest counter-camp: Jensen-Kelly-Pedersen (2023, JF) find 55-61% of factors do replicate and are strengthened by the number of correlated factors; Chen-Zimmermann show only ~half of the decay is publication bias. The replication rate is genuinely unresolved and methodology-dependent - so do not treat either number as settled. But this dispute is about equity asset-pricing factors; it rescues the structural survivors, not dead price patterns.

Capacity is the retail edge (the hard proof)

The single most useful result for a solo trader: Jacobs & Müller (2020, JFE) studied 241 anomalies across 39 markets and found the US is the only country with a reliable post-publication decline. Anomalies survive internationally, where arbitrage is costly.

The implication inverts the usual worry. Edges do not die everywhere at once - they die where cheap, deep arbitrage capital removes them. Small and illiquid markets keep edges that are uneconomic to harvest at institutional scale, not hidden. For a trader with small capital, the advantage is capacity, not information: you can hold things that are too small for a fund to bother with.

What survives, and why

Carry (Koijen-Moskowitz-Pedersen-Vrugt, 2018 JFE): a general premium across equities, bonds, FX, commodities, credit. Within-class Sharpe ~0.7, diversified ~1.1, near-zero cross-class correlation. It persists because its three worst drawdowns land on global recessions - you are paid for that. Crypto funding carry is one sleeve of this, not a standalone anomaly.

Variance risk premium / short-vol (Heston-Todorov; Bollerslev-Todorov): negative in 19 of 20 futures; roughly half is direct compensation for jump/disaster risk (left tail 59.8% vs right 10.0%). Profitable most years, structurally short the tail - exactly Botty's VRP-harvest shape.

Trend / time-series momentum (Moskowitz-Ooi-Pedersen, 2012 JFE): diversified Sharpe >1, but that is in-sample 1965-2009. Treat magnitude as a historical upper bound.

Breadth, not depth - the strongest lever

The number that retroactively explains Botty's entire BTC-timing result:

  • AQR 'Trends Everywhere': diversified multi-market trend gross Sharpe ~1.60 vs ~0.35 for the median single instrument. A factor of ~4.5.
  • Hurst-Ooi-Pedersen (century evidence): net-of-fees Sharpe ~1.0, positive in every decade 1903-2012.
  • Single-asset trend since 1800: ~0.3 vs diversified ~0.7.

Botty's single-asset BTC-Donchian is the 0.35 version. It was never a bad signal - it was the same trend signal with its noise left un-diversified. The evidence-strongest direction is not a better BTC entry; it is the same slow trend logic spread across many uncorrelated markets.

Post-2009 caveat, stated honestly: fast trend collapsed (EWM-5-20 Sharpe 0.84 to 0.12), slow trend held (EWM-50-200 0.70 to 0.40). The deployable version is slow, diversified, cost-aware - not fast.

Do backtests still make sense?

Yes - as a falsifier, not a discoverer. With enough trials, noise guarantees a great-looking backtest; Botty's own megasweep PBO ~0.76 is that fact measured. The megasweeps were not wasted - they were a correctly-working falsification machine, just pointed at a hypothesis space (chart patterns) with no economic payer.

The fix is order-of-operations. First the economic question - who is paying me, and for what risk? Then the backtest, to verify and size. That order is precisely what divides the survivors from the dead.

Honest limits of this research

  • Domain mismatch: nearly all primary evidence is equity / cross-asset. None of it validates crypto TA - fully consistent with Botty's own OOS-death finding, but it means the survivors are argued by analogy to their asset-class cousins, not by crypto-specific proof.
  • No small-perp number: the capacity argument is supported in principle, but there is no ready net-of-cost, net-of-delisting yield figure for small crypto perps. That has to be benched in-house before any capital moves.
  • Backtest-overfitting pillar (PBO/DSR/Lopez de Prado) came through only indirectly via the multiple-testing literature, not as a directly verified source in this pass.