Learning Objectives
- Build split-aware preprocessing pipelines that produce stable, auditable inputs for label and feature computation.
- Define execution-consistent labels, including fixed-horizon and event-style constructions, and diagnose overlap, resolution behavior, and implied trading intensity.
- Evaluate feature-label bundles fold by fold using appropriate diagnostics for continuous and discrete targets, including stability, shape, and feasibility.
- Screen candidates for implementation feasibility using turnover, break-even cost, and liquidity or capacity checks.
- Account for search bias by defining searched sets, separating exploration from confirmation, and applying appropriate multiple-testing adjustments to fold-level summaries.
- Use mechanism plausibility checks to distinguish potentially stable signal channels from confounded proxies, timing artifacts, and aggregation effects.
Data preprocessing and encodings
Label engineering
Univariate feature–label evaluation
Search accounting and multiple testing
From correlation to causality
Summary
Related Case Studies
See where these chapter concepts get applied in end-to-end trading workflows.
ETF Cross-Asset Exposures
All six model families compared across 100 ETFs spanning 9 asset classes
Crypto Perpetuals Funding
Alternative data and non-standard frequencies in 24/7 crypto markets
NASDAQ-100 Microstructure
Intraday microstructure signals across 114 stocks at 15-minute frequency
S&P 500 Equity + Option Analytics
Combining options-derived features with equity data for multi-source prediction
US Firm Characteristics
Classic factor investing with ML on monthly fundamental data
FX Spot Pairs
Momentum and carry factors in the world's most liquid market
CME Futures
Carry signals across 30 products — data quality as the critical variable
S&P 500 Options (Straddles)
Direct options trading and why equity-style cost models fail for options
US Equities Panel
Large-scale cross-sectional prediction across 3,200 stocks with 16 walk-forward folds
01 Data Quality Diagnostics
02 Preprocessing Pipeline
03 Label Methods
04 Maximum Favorable Adverse Excursion
05 Signal Evaluation
06 Ic Inference
07 Multiple Testing
08 Causal Sanity Checks
10 Ml4T Library Ecosystem
18 primer topics providing foundational concepts for this chapter.
Block Bootstrap and Permutation Testing for Dependent Data
Once observations have an order, valid resampling must preserve more than the one-period distribution.
Causality, Confounding, and Why Good Signals Can Be Misleading
A predictive relationship can be real in the data and still fail as an explanation of what would happen under intervention.
Coverage-Aware Evaluation and Event-Time Alignment for Text Signals
A text model is not useful because it predicts labels accurately. It is useful only if its signal is available when you trade, on enough names, at the horizon that matters.
From Information Coefficient to Information Ratio
"It takes only a modest amount of skill to win as long as that skill is deployed frequently and across a large number of stocks." -- Grinold and Kahn (1999)
HAC Standard Errors and Robust Inference
When forecast errors have memory, the coefficient may stay the same while the t-statistic should not.
Hypothesis Testing and P-Values
How hypothesis tests turn noisy evidence into a structured decision, and how to read p-values without treating them as proof.
Label Overlap: Why Your Sample Is Smaller Than You Think
When labels share future price paths, nominal sample counts exaggerate the evidence available for inference — often by an order of magnitude. Diagnosing overlap before interpreting results is not optional; it is the difference between a credible signal evaluation and a statistical illusion.
Momentum and Mean Reversion
How return predictability changes with horizon, how cross-sectional and time-series momentum differ, and why the same signal that works for months can fail violently in a rebound.
Multiple Testing and the Researcher’s Trap
Why searching many ideas makes false discoveries and overstated winners inevitable.
Multiple Testing in Factor Research: The Search Tax on Discovery
Every variant you try without recording it borrows from the credibility of the winner. The statistical correction is straightforward; the organizational discipline of tracking what you searched is harder and more important.
Multiple Testing, Replication, and the Factor Zoo After the Replication Wars
The factor zoo is not just a story about too many predictors. It is a story about search, construction choices, and the gap between a published anomaly and an investable factor.
Point-in-Time Data and Decision-Time Correctness
A value is usable only if it was actually knowable when the strategy had to decide.
Reading the Information Coefficient: Stability, ICIR, and Horizon Decay
IC is a learnability screen for continuous labels, not a compressed backtest.
Simple Returns vs Log Returns
One aggregates exactly across assets, the other aggregates exactly across time. Most mistakes come from asking one definition to do both jobs.
The Information Coefficient
How cross-sectional rank correlation measures signal quality, and why a small edge can still matter when it is applied repeatedly.
Trading Costs: Spread, Slippage, and Market Impact
How execution costs arise, why the components are different, and how turnover turns a predictive signal into a net strategy.
Volatility: Realized, Implied, and Why It Clusters
Volatility is not one number but a family of related objects: what happened, what the options market prices, and what a model forecasts next.
Walk-Forward Validation for Time Series
Why model evaluation must preserve temporal order, and how expanding or rolling splits approximate live deployment.
66 references cited in this chapter.
Whitney K. Newey and Kenneth D. West (1986)
Narasimhan Jegadeesh and Sheridan Titman (1993) — The Journal of Finance · 11065 citations
This seminal paper establishes the existence of 'Momentum' in stock prices, showing that buying past 6-month winners and selling losers generates ~1% monthly excess returns over medium horizons (3-12 months).
Yoav Benjamini and Yosef Hochberg (1995) — Journal of the Royal Statistical Society. Series B (Methodological)
Ardian Harri and B. Wade Brorsen (1998) — SSRN Electronic Journal
This paper explains why long-horizon/overlapping-return regressions create moving-average errors that break standard inference, and it provides practical guidance on which estimator (GLS, MLE, NW, non-overlap) is appropriate in each common finance use case.
Halbert White (2000) — Econometrica · 1769 citations
This paper introduces a 'Reality Check' procedure to test whether the best model found during a specification search has genuine predictive power over a benchmark, accounting for data snooping biases.
Active portfolio management: A quantitative approach for providing superior returns and controlling risk
Richard C.. Grinold and Ronald N.. Kahn (2000) — McGraw-Hill · 222 citations
Clifford S. Asness et al. (2013) — The Journal of Finance
Value and Momentum strategies work consistently across eight diverse asset classes and markets, and their negative correlation allows a combined strategy to achieve a Sharpe ratio of 1.45, significantly outperforming either strategy alone.
David H. Bailey and Marcos Lopez de Prado (2014) · 110 citations
The paper introduces the Deflated Sharpe Ratio (DSR), a statistic that adjusts performance metrics for the probability of backtest overfitting caused by multiple testing and non-normal returns.
Marcos Lopez de Prado (2015)
This paper argues that most published backtests in empirical finance are statistically unreliable because they ignore multiple testing and other sources of performance inflation, and it proposes practical corrections (Deflated Sharpe Ratio and Probability of Backtest Overfitting).
David H. Bailey et al. (2015) · 119 citations
This paper proposes a practical, model-free way to estimate how likely your “best” backtest is a false positive caused by trying many strategy variations on the same data.
...and the Cross-Section of Expected Returns
Campbell R. Harvey et al. (2016) — Review of Financial Studies · 1838 citations
Due to extensive data mining in asset pricing ('the factor zoo'), the standard t-statistic threshold of 2.0 is insufficient; this paper mathematically demonstrates that a t-statistic > 3.0 is required to establish true significance.
Does Academic Research Destroy Stock Return Predictability?
R. David McLean and Jeffrey Pontiff (2016) — Journal of Finance · 851 citations
Stock return predictors published in academic journals lose 58% of their performance after publication due to a combination of statistical bias (data mining) and arbitrage trading.
Kent Daniel and Tobias J. Moskowitz (2016) — Journal of Financial Economics · 2 citations
Momentum strategies suffer predictable crashes during bear market rebounds because past losers behave like high-beta call options; a dynamic weighting strategy based on market volatility can double the Sharpe ratio.
Robert Novy-Marx and Mihail Velikov (2016) — The Review of Financial Studies · 513 citations
The paper shows that many well-known equity anomalies look attractive before costs but lose most (or all) of their profitability after realistic trading costs—unless you redesign the strategy using a simple buy/hold “sS-rule” that sharply reduces turnover.
Advances in Financial Machine Learning
Marcos Lopez de Prado (2018) — John Wiley & Sons · 106 citations
Andrea Frazzini et al. (2018) · 112 citations
Using $1.7T of live institutional executions across 21 developed equity markets (1998–2016), the paper measures real-world implementation shortfall/price impact and finds costs for a patient large trader are far smaller than standard TAQ-based academic estimates, and are best described by a concave (≈ square-root) impact function in trade size.
Judea Pearl (2019) — Communications of the ACM · 620 citations
Pearl argues that moving from pattern recognition to causal models (especially interventions and counterfactuals) is necessary to make ML systems robust, explainable, and capable of “what if?” reasoning.
Kewei Hou et al. (2020) — The Review of Financial Studies · 759 citations
This paper replicates 452 published financial anomalies and finds that the majority fail to hold up to current empirical finance standards, particularly when controlling for microcaps and multiple testing, suggesting capital markets are more efficient than previously recognized.
Carlos Cinelli and Chad Hazlett (2020) — Journal of the Royal Statistical Society. Series B (Statistical Methodology) · 837 citations
Feng Zhang et al. (2020)
The paper explains why Information Coefficient (IC) is usually tiny and very noisy in realistic stock-selection settings, and proposes two concrete hypothesis-testing procedures to monitor whether a model’s IC is deteriorating over time.
Andrew Y. Chen and Tom Zimmermann (2021) · 185 citations
This paper provides an open-source dataset and code demonstrating that nearly 100% of published cross-sectional stock return predictability results can be successfully reproduced, challenging claims of widespread non-replicability and p-hacking.
Stefano Giglio et al. (2021) — The Review of Financial Studies
This paper develops a robust framework for multiple hypothesis testing of alphas in linear asset pricing models, controlling for data snooping, omitted factors, and missing data, and demonstrates its superior performance in hedge fund evaluation.
Ilia Zaznov et al. (2022) — Mathematics
A critical review and reproduction study demonstrating that while deep learning models achieve >80% accuracy on LOB benchmarks, they fail to generate profit when realistic bid-ask spreads and transaction costs are applied.
Is There a Replication Crisis in Finance?
Theis Ingerslev Jensen et al. (2022) · 285 citations
This paper argues that the perceived replication crisis in financial economics is overstated, demonstrating through a novel hierarchical Bayesian model and global data that most asset pricing factors are replicable, cluster into 13 themes, and show external validity.
Mihir Tirodkar (2025) — The Journal of Portfolio Management
Using forward-looking option-implied returns rather than noisy realized returns, this paper demonstrates that momentum is not a priced risk factor but rather a dynamic strategy that generates negative expected returns during market crises.
Marcos Lopez de Prado and Vincent Zoonekynd (2025)
This paper argues that many failed factor strategies are not just p-hacked but causally misspecified ('Factor Mirages'), and demonstrates how standard regression techniques involving colliders can flip the sign of risk premia.
Paul Glasserman et al. (2025)
Using 2.4M Reuters articles (1996–2022), the paper shows that persistent firm-level news-topic exposures and opposite intraday vs overnight price responses to the same topics explain a large share of the long-run U.S. “overnight drift” (overnight gains vs flat/negative intraday returns) and help forecast which stocks will do best overnight and worst intraday.
Does Peer-Reviewed Research Help Predict Stock Returns?
Andrew Y Chen et al.
Peer-reviewed cross-sectional stock return predictors deliver only a ~2 bps/month out-of-sample advantage versus a very naive data-mining procedure on 29,000 accounting ratios, and “risk-based” theoretical justifications do not improve (and may worsen) post-sample robustness.
Tom Fawcett — Pattern Recognition Letters
This paper explains how to use ROC curves/AUC correctly to evaluate and select classifiers—especially under class imbalance and changing misclassification costs—while avoiding common pitfalls like score miscalibration and tied-score artifacts.
Robert Novy-Marx
The paper shows that backtests of strategies built by combining and signing multiple signals are severely overfit—so conventional t-stat thresholds (e.g., 1.96) dramatically overstate statistical significance, often requiring critical values as high as ~4–7 instead.
Eugene F. Fama and Kenneth R. French (2008) — The Journal of Finance
Fama and French demonstrate that many popular anomalies (asset growth, profitability) disappear in large-cap stocks and are driven by illiquid microcaps, while momentum and net stock issues remain robust across all size groups.
On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation
Gavin C Cawley and Nicola L C Talbot (2010) · 2191 citations
This paper demonstrates that the variance of model selection criteria can lead to overfitting hyperparameters, resulting in poor generalization and an optimistic bias in machine learning performance evaluations if proper nested validation protocols are not used.
Mark Cummins and Andrea Bucca (2012) — Quantitative Finance
The paper tests an Ornstein–Uhlenbeck (OU) mean-reversion statistical-arbitrage model across 861 oil-related futures spreads and—after controlling for data-snooping—finds many profitable spread trades (often Sharpe > 2), except during the 2008 crisis regime shift when profitability collapses.
David H. Bailey et al. (2014) — Notices of the American Mathematical Society · 150 citations
The paper shows mathematically and via simulations how trying many strategy variations can manufacture high in-sample Sharpe ratios that collapse (or even turn negative) out-of-sample, and it proposes a “minimum backtest length” requirement tied to the number of trials.
David Colquhoun (2014) — Royal Society Open Science
The paper shows (with simple probability arguments and t-test simulations) that treating p = 0.05 as “a discovery” typically implies a false discovery rate around 30% or higher—often much higher in low-power studies—so p-values are routinely over-interpreted.
Nikolay Gospodinov et al. (2014) — The Review of Financial Studies
This paper demonstrates that in misspecified asset-pricing models, irrelevant risk factors (uncorrelated with test asset returns) can lead to erroneous conclusions about factor pricing, and proposes a robust model selection procedure to address this issue.
Online tools for demonstration of backtest overfitting
David H Bailey et al. (2015) — Available at SSRN 2597421
This paper introduces two free online simulators (BODT and TMST) that show how optimizing many strategy variants on the same backtest can manufacture high in-sample Sharpe ratios that collapse out-of-sample—and why metrics like the Deflated Sharpe Ratio are needed to correct for multiple testing.
Zura Kakushadze et al. (2015) · 10 citations
A catalog of 101 explicit, executable quantitative trading alpha formulas with empirical analysis showing returns scale with volatility (exponent ~0.76) but are independent of turnover.
Jeremiah Green et al. (2017) — The Review of Financial Studies
By simultaneously testing 94 firm characteristics, the authors find that only 12 provide independent return information for non-microcap stocks, and this predictability effectively vanished after 2003 due to increased market efficiency.
Campbell R. Harvey (2017) — The Journal of Finance · 286 citations
This paper discusses the overuse and misuse of p-values in financial economics research, highlighting issues like p-hacking and publication bias, and suggests adopting a Bayesian perspective for more robust and transparent research practices.
Campbell R. Harvey and Yan Liu (2019) · 70 citations
A comprehensive meta-analysis of over 400 published investment factors, arguing that most are false positives resulting from p-hacking and multiple testing failures.
Nicolas Huck (2019) — European Journal of Operational Research
A large-scale backtest (1993–2015) shows that ML models can generate short-horizon long–short signals on US large-caps, but performance collapses after costs and risk adjustment—especially after 2010—and adding hundreds of features often hurts rather than helps.
Guanhao Feng et al. (2020) — The Journal of Finance · 490 citations
The authors propose a Double-Selection LASSO methodology to evaluate new asset pricing factors against hundreds of existing ones, finding that while most new factors are redundant, profitability and investment factors provide significant marginal value.
Alik Sokolov et al. (2023)
The paper proposes an end-to-end, LLM-assisted pipeline that generates causal DAGs from feature definitions (not just correlations) and then empirically prunes/validates edges with do-calculus—demonstrated on equity factor investing and credit-spread drivers.
Amit Goyal et al. (2024) — The Review of Financial Studies
This paper re-evaluates the predictive power of 29 recently published variables, along with 17 original variables, for forecasting the equity premium, finding that many have lost their in-sample and out-of-sample significance when tested with data extended to the end of 2021.
Andrew Y. Chen (2024)
Contradicting the 'replication crisis' narrative, this paper demonstrates that at least 75-91% of published cross-sectional predictors are statistically valid, arguing that previous claims of widespread false discoveries stem from misinterpreting statistical insignificance as falsity.
Jacques Joubert et al. (2024) — The Journal of Portfolio Management
A practitioner-focused guide to designing more reliable backtests by choosing the right backtest type, avoiding common simulation biases, and correcting Sharpe-ratio inference for multiple strategy trials.
Eve Richardson et al. (2024) — Patterns
This paper shows that ROC-AUC is invariant to class imbalance (when class imbalance does not change the per-class score distributions), whereas PR-AUC changes strongly with imbalance, so ROC-AUC is suitable for comparing models across datasets with different base rates.
Matthew B. McDermott et al. (2024) — Advances in Neural Information Processing Systems
This paper shows that the common rule-of-thumb “use AUPRC instead of AUROC under class imbalance” is generally false, and that AUPRC can even worsen subgroup fairness by rewarding improvements concentrated in high-score/high-prevalence regions.
Marcos Lopez de Prado et al. (2024)
This paper argues that mainstream (associational) factor model specification practices systematically select misspecified models—especially by mis-handling confounders and colliders—causing disappointing live factor performance and even systematic losses, and it calls for rebuilding factor research around causal graphs and do-calculus.
Nicolas Harvie et al. (2024)
This paper demonstrates that commonly used heteroskedasticity- and autocorrelation-consistent (HAC) estimators can lead to inflated t-statistics and invalid inference in asset pricing, and proposes a simulation-based inference procedure called SHARFS to address this issue.
Matias D. Cattaneo et al. — The Review of Economics and Statistics
This paper formalizes portfolio sorting as a nonparametric estimator, proving that the standard practice of using 5 or 10 fixed portfolios is suboptimal and providing a data-driven method to select the optimal number of portfolios (often 50+).
J. Scott Armstrong and Fred Collopy — International Journal of Forecasting
The paper empirically compares common forecast error metrics and shows that RMSE is too unreliable for comparing methods across many time series, recommending relative/robust alternatives like GMRAE, MdRAE, and MdAPE depending on the task and sample size.
Tilmann Gneiting — Journal of the American Statistical Association
Point forecasts can be ranked incorrectly if the error metric (scoring function) is not aligned with what the forecast is supposed to represent (mean, median, quantile, etc.), so you must specify the loss ex ante or evaluate only with scoring functions consistent for a named functional.
Ryan Sullivan et al. — The Journal of Finance
This paper applies White's Reality Check bootstrap methodology to rigorously evaluate thousands of technical trading rules, finding that while they show significant in-sample performance even after data-snooping adjustment, this performance largely disappears out-of-sample, particularly in liquid futures markets.
Philip Hans Franses — International Journal of Forecasting
The paper argues that MASE (Mean Absolute Scaled Error) is a forecast accuracy measure that “plays nicely” with Diebold–Mariano equal-accuracy tests because it preserves the moment conditions needed for asymptotic normality, unlike percentage-error based measures.
Antonios Antoniou et al. — Journal of Banking & Finance
This paper demonstrates that momentum strategies are highly profitable in UK, German, and French markets and argues that these profits are driven by risk factors related to the business cycle rather than purely behavioral biases.
Victor Chernozhukov et al. (2018) — The Econometrics Journal · 3298 citations
This paper shows how to get valid (root-N, asymptotically normal) inference for causal/structural parameters when nuisance components are learned with high-dimensional ML, by combining Neyman-orthogonal scores with cross-fitting.
Daniel Fryer et al. (2021) — IEEE Access
This paper argues that using Shapley-value feature attributions as a feature selection rule (e.g., “pick the top-k Shapley features”) can be fundamentally misaligned with feature selection goals because the Shapley axioms encourage model-averaged “fair credit,” not “choose the best subset.”
Jules H. van Binsbergen et al. (2024)
This corrigendum fixes a look-ahead/data-leakage issue in an earnings-forecasting ML setup and shows that ML still beats analysts on forecast accuracy, while the return-predictability linked to analyst optimism remains but is quantitatively weaker.
Jules H. van Binsbergen et al. (2025)
This paper rebuts claims that a look-ahead bias invalidates the original findings of Binsbergen et al. (2023), demonstrating that machine learning models still outperform analysts and linear models in earnings forecasting and return predictability even after correcting for the bias.
Rob J. Hyndman and Anne B. Koehler — International Journal of Forecasting
The paper shows that many popular forecast-accuracy metrics (especially MAPE/sMAPE and several M-competition/M3 measures) break down with zeros, negatives, or flat stretches, and proposes MASE as a robust, interpretable default for comparing forecast accuracy across many time series.
Ming Liu et al. — Journal of International Money and Finance
The 52-week high strategy generates robust returns across 18 of 20 international markets, distinct from traditional momentum, but transaction costs render it statistically insignificant in most regions.
Ben R. Marshall et al. — Journal of Banking & Finance
A rigorous empirical test of 28 candlestick patterns on DJIA stocks using a bootstrap methodology finds no evidence of profitability, supporting market efficiency in large-cap US equities.
Keywan Christian Rasekhschaffe and Robert C. Jones — Financial Analysts Journal
This paper demonstrates that combining forecasts from multiple machine learning algorithms and distinct training windows (recent, seasonal, and 'hedge') significantly outperforms linear models and reduces overfitting in stock selection.