The ML Pipeline
Learning Objectives
- Choose between regression and classification formulations based on how predictions will be translated into trading decisions
- Fit leakage-safe regularized linear models, including Ridge, LASSO, Elastic Net, and logistic regression, using point-in-time preprocessing and standardization
- Tune and evaluate linear models with walk-forward validation, temporal buffers, and, when needed, nested cross-validation to reduce selection bias
- Interpret model behavior with SHAP-based diagnostics to assess feature importance, economic plausibility, and stability across refits
- Construct and evaluate conformal prediction intervals or prediction sets, and monitor where coverage degrades under non-stationary market conditions
- Use cross-case-study evidence to judge when linear models provide a strong baseline and when weak linear signal motivates more flexible models
From inference to prediction
This section argues that the shift from econometric inference to predictive modeling changes which estimator properties matter: unbiasedness (the Gauss-Markov criterion) gives way to minimizing total prediction error via the bias-variance tradeoff. It develops the case through the two-cultures framework (Breiman 2001), showing how high dimensionality, multicollinearity, and low signal-to-noise ratios in financial features make unconstrained OLS unsuitable for forecasting. The section introduces Ridge, LASSO, and Elastic Net as principled responses that encode different structural priors about how signal is distributed across features, and maps the label families from Chapter 7 to regression and classification tasks.
Regularized regression
The section presents Ridge, LASSO, and Elastic Net with their formal objectives, geometric intuitions, and distinctive behaviors on correlated financial features. It covers the full practical machinery required for deployment: standardization with leakage-safe protocols, hyperparameter optimization via Optuna with nested cross-validation to guard against validation overfitting, loss function choice (MSE, MAE, Huber, quantile), evaluation metrics (IC, ICIR, RMSE), and sample weighting for overlapping labels and recency adaptation. Empirical results on the ETF case study demonstrate that Ridge achieves a 1.5x ICIR improvement over OLS at optimal regularization, while LASSO degrades performance because the ETF signal is diffusely distributed across correlated features.
Predicting direction with logistic regression
This section develops logistic regression as the regularized classification baseline for discrete trading decisions, covering binary and multinomial formulations, probability calibration, and the conversion from predicted probabilities to trading signals via threshold-based, probability-weighted, and rank-based methods. It explains why regularization is even more critical for classification than regression (the maximum likelihood objective diverges under near-perfect separability) and addresses class imbalance through inverse-frequency weighting. The section provides practical guidance on when classification outperforms regression (binary trading decisions, noisy return targets) versus when it underperforms (asymmetric payoffs, magnitude-dependent position sizing).
Interpreting models with SHAP
The section establishes SHAP as the primary interpretability framework, grounded in Shapley values from cooperative game theory, which uniquely satisfy efficiency, symmetry, null-player, and linearity axioms. It develops a four-layer economic narrative protocol — sign consistency, magnitude plausibility, stability across walk-forward windows, and regime-conditional analysis — that transforms SHAP from a visualization tool into a continuous diagnostic for distinguishing genuine signal from overfitting. The section demonstrates both global feature importance and local waterfall explanations on the ETF case study, and addresses limitations including causal misinterpretation and impossible coalitions from marginalizing correlated features.
Quantifying predictive uncertainty
This section introduces conformal prediction as a distribution-free framework for constructing prediction intervals that wrap around any base model without distributional assumptions. It progresses from split-conformal prediction (fixed-width intervals with finite-sample marginal coverage guarantees) through Conformalized Quantile Regression (adaptive-width intervals) to Adaptive Conformal Inference (online coverage correction for non-stationary data). Empirical results on the ETF case study show that CQR+ACI progressively closes the conditional coverage gap during high-volatility periods (from 82.3% to 88.1% for a 90% target), and the section previews how interval width maps to position sizing in Chapter 19.
Case study insights
Running the same regularized pipeline across all nine case studies reveals that linear signal availability varies dramatically by asset class and market structure: delta-hedged options and ETFs show the strongest ICs, while CME futures and crypto are near zero on primary labels. The section demonstrates that label preprocessing (winsorization, horizon selection) often dominates model selection in its effect on IC, that Ridge wins the clear majority of primary-label comparisons due to the pervasive correlation structure of financial features, and that classification can outperform raw-return regression on noisy targets. A pedagogical backtest shows that Ridge's IC advantage over simple momentum is eroded by higher turnover, establishing the IC-cost tradeoff as the central tension for Chapters 16-19.
Summary
Related Case Studies
See where these chapter concepts get applied in end-to-end trading workflows.
ETF Cross-Asset Exposures
All six model families compared across 100 ETFs spanning 9 asset classes
Crypto Perpetuals Funding
Alternative data and non-standard frequencies in 24/7 crypto markets
NASDAQ-100 Microstructure
Intraday microstructure signals across 114 stocks at 15-minute frequency
S&P 500 Equity + Option Analytics
Combining options-derived features with equity data for multi-source prediction
US Firm Characteristics
Classic factor investing with ML on monthly fundamental data
FX Spot Pairs
Momentum and carry factors in the world's most liquid market
CME Futures
Carry signals across 30 products — data quality as the critical variable
S&P 500 Options (Straddles)
Direct options trading and why equity-style cost models fail for options
US Equities Panel
Large-scale cross-sectional prediction across 3,200 stocks with 16 walk-forward folds
01 Ols Inference
02 Regularization Paths
03 Logistic Classification
04 Nested Cv Hpo
05 Shap Analysis
06 Conformal Prediction
07 Case Study Insights
08 Ml Backtest Intro
15 primer topics providing foundational concepts for this chapter.
Bayesian Hyperparameter Optimization Under Temporal Dependence
Hyperparameter search is part of the statistical design, not a software convenience layer.
Classical Statistical Tests as Linear Models
OLS, t-tests, ANOVA, and Pearson correlation are not separate islands; for the cases covered here, they are linear models with different design matrices and different coefficient restrictions.
Conformal Prediction in Finance: Coverage, Exchangeability, and Drift
Conformal prediction is attractive in finance because it gives finite-sample coverage without distributional assumptions, but its guarantee is only as good as the exchangeability you are willing to believe.
Estimation Error and the Markowitz Curse
Mean-variance optimization is not fragile because the quadratic program is hard; it is fragile because the optimizer is asked to invert noisy beliefs about returns and covariances.
Hypothesis Testing and P-Values
How hypothesis tests turn noisy evidence into a structured decision, and how to read p-values without treating them as proof.
Label Overlap: Why Your Sample Is Smaller Than You Think
When labels share future price paths, nominal sample counts exaggerate the evidence available for inference — often by an order of magnitude. Diagnosing overlap before interpreting results is not optional; it is the difference between a credible signal evaluation and a statistical illusion.
Loss Functions, Error Metrics, and What They Hide
A model is trained to optimize one quantity, selected on another, and traded on a third. Most confusion in predictive modeling starts when those three layers are blurred together.
Proper Scoring Rules
A probability forecast is only as useful as the score that judges it: the right score rewards honest beliefs, while the wrong one rewards distortion.
Proper Scoring Rules for Financial Event Forecasts
How to evaluate agent-generated probabilities so that honesty, calibration, and useful discrimination all show up in the score.
Regularization Geometry: How Ridge, LASSO, and Elastic Net Actually Work
Regularization helps not by fitting the training sample better, but by refusing to trust unstable coefficient estimates — and the SVD of the feature matrix reveals exactly which directions it distrusts and why.
Selection Bias in Model Tuning: Why Your Best Validation Score Lies
Even with perfect chronological splits and no data leakage, repeated hyperparameter search overfits the validation set — and the winning score systematically overstates the performance you should expect out of sample.
The Bias-Variance Tradeoff
Why a model that is deliberately a little wrong can generalize better than one that fits the past too closely.
Uncertainty as a Feature
Why two identical forecasts can imply very different decisions once you look at what the model does not know.
Uncertainty Estimation and Calibration for Deep Time-Series Models
A forecasting model is not uncertainty-aware because it emits a variance. It is uncertainty-aware only if that variance tracks future error under the validation protocol you actually trade.
Walk-Forward Validation for Time Series
Why model evaluation must preserve temporal order, and how expanding or rolling splits approximate live deployment.
143 references cited in this chapter.
Portfolio selection
Harry Markowitz (1952) — The journal of finance · 5471 citations
Markowitz formalizes modern portfolio theory by showing that portfolio choice should trade off expected return against variance (risk), producing an efficient frontier of optimal diversified portfolios.
Harold William Kuhn et al. (1953) — Princeton University Press · 1998 citations
Peter J. Huber (1964) — The Annals of Mathematical Statistics · 2194 citations
This paper introduces a theory for robust estimation of a location parameter, particularly for contaminated normal distributions, and identifies estimators that are asymptotically most robust among translation invariant estimators, offering a balance between the sample mean and sample median.
Edward O. Thorp (1975) — Elsevier · 216 citations
Thorp explains why maximizing expected log-wealth (the Kelly criterion) is a principled, growth-optimal portfolio rule, clarifies common objections (especially Samuelson’s), and illustrates real-world implementation via convertible-hedge arbitrage performance.
Louis M. Rotando and Edward O. Thorp (1992) — The American Mathematical Monthly · 96 citations
This paper adapts the Kelly Criterion (optimal betting sizing for geometric growth) to the U.S. stock market, finding that historically the optimal leverage for the S&P 500 was 117%.
Robert Tibshirani (1996) — Journal of the Royal Statistical Society: Series B (Methodological) · 50893 citations
This paper introduces the 'lasso,' a new linear model estimation technique that minimizes residual sum of squares with a constraint on the sum of the absolute values of the coefficients, effectively shrinking some coefficients to zero for model interpretability and improved prediction accuracy.
Statistical Modeling: The Two Cultures
Leo Breiman (2001) · 3769 citations
This paper argues that the statistical community's over-reliance on data models has led to irrelevant theory and missed opportunities, advocating for a shift towards algorithmic modeling which focuses on predictive accuracy and can handle complex, unknown data mechanisms.
Harris Papadopoulos et al. (2002) — Springer-Verlag · 605 citations
This paper introduces Inductive Confidence Machines (ICM) for regression, a computationally efficient method that provides prediction intervals with associated confidence levels, addressing the inefficiency of previous transductive methods.
International Asset Allocation With Regime Shifts
Andrew Ang and Geert Bekaert (2002) — Review of Financial Studies · 1567 citations
Despite correlations rising in bear markets, international diversification remains economically valuable, particularly when investors can switch into cash (risk-free assets) during high-volatility regimes.
Olivier Ledoit and Michael Wolf (2003) — Journal of Empirical Finance · 1639 citations
The paper introduces an “optimal shrinkage” estimator that blends the sample covariance matrix with a single-index (market) covariance model, producing more stable covariance estimates and materially lower out-of-sample minimum-variance portfolio risk on NYSE/AMEX stocks (1972–1995).
Alexandru Niculescu-Mizil and Rich Caruana (2005) — Association for Computing Machinery · 1862 citations
Hui Zou and Trevor Hastie (2005) — Journal of the Royal Statistical Society Series B: Statistical Methodology · 9404 citations
This paper introduces the elastic net, a new regularization and variable selection method that combines the strengths of both lasso and ridge regression, often outperforming the lasso, especially when dealing with highly correlated predictors or when the number of predictors is much larger than the number of observations.
Chris Whitrow (2007) — Journal of the Royal Statistical Society Series C: Applied Statistics · 13 citations
Whitrow shows how to compute Kelly-optimal stake allocations across many simultaneous bets using fast Monte Carlo stochastic-gradient algorithms, avoiding inaccurate closed-form approximations when edges are not tiny or the bet set is large.
Sébastien Maillard et al. (2008) — SSRN Electronic Journal · 401 citations
The paper formalizes “equal risk contribution” (ERC) / risk parity portfolios, proves key properties (including a volatility ordering between minimum-variance and 1/n), and shows in rolling backtests that ERC can deliver a practical diversification/turnover trade-off versus minimum-variance and equal-weighting.
Trevor Hastie et al. (2009) — Springer-Verlag · 20470 citations
The paper shows that, for very small antimicrobial-peptide datasets (n≈101), model selection should prioritize statistical stability (via learning curves, bootstrapping, and permutation tests) rather than one-off peak metrics, and that GA-based feature selection and stacked autoencoders yield competitive, statistically significant SVR predictors.
On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation
Gavin C Cawley and Nicola L C Talbot (2010) · 2191 citations
This paper demonstrates that the variance of model selection criteria can lead to overfitting hyperparameters, resulting in poor generalization and an optimistic bias in machine learning performance evaluations if proper nested validation protocols are not used.
Understanding Risk Parity
Brian Hurst (2010)
The paper argues that “60/40” portfolios are not truly diversified because equity risk dominates, and shows that risk-parity (risk-balanced) allocations across equities, bonds, and commodities historically delivered higher risk-adjusted returns and smaller drawdowns at comparable volatility.
Paolo Laureti et al. (2010) — Quantitative Finance · 21 citations
The paper analyzes Kelly (log-growth) portfolio optimization in a simple lognormal-return model, deriving closed-form approximations, showing when Kelly aligns with Markowitz efficient frontiers, and explaining why Kelly portfolios often under-diversify (“condense”) onto few assets.
J. L. Kelly (2011) — WORLD SCIENTIFIC
Kelly links Shannon’s channel information rate to optimal long-run (log-utility) capital growth: the maximum exponential bankroll growth achievable from noisy signals equals the channel’s mutual information.
Edward O. Thorp (2011) — WORLD SCIENTIFIC · 42 citations
Thorp explains what the Kelly criterion really optimizes (expected log-wealth / long-run growth), and why naïvely applying “edge/odds” or single-asset Kelly fractions can be dangerously wrong in real investing.
Jun Liu and Allan Timmermann (2013) — The Review of Financial Studies · 77 citations
The paper derives utility-maximizing (often non–delta-neutral) trading rules for convergence/pairs trades when the spread is cointegrated, showing that standard 1/-1 long–short arbitrage can be materially suboptimal—especially when convergence is “one-shot” (stopped) or horizons are short or random.
Vasily Nekrasov (2014) · 26 citations
The paper shows how to approximate and numerically compute Kelly-optimal portfolio weights for many correlated assets using only first and second moments, and proposes a GPU-friendly Monte Carlo “groping” algorithm for fractional Kelly portfolios under realistic constraints.
Marcos Lopez de Prado (2016) · 40 citations
The paper proposes Hierarchical Risk Parity (HRP), a clustering-based portfolio construction method that avoids covariance-matrix inversion and delivers materially lower out-of-sample variance than Markowitz-style minimum-variance and simple risk parity in Monte Carlo tests.
Thomas Raffinot (2016) · 100 citations
The paper proposes a simple portfolio construction method that allocates capital using hierarchical clustering of asset correlations (including DBHT/PMFG) and shows, across three datasets, that these portfolios are more diversified and deliver statistically superior out-of-sample risk-adjusted performance than standard risk-based optimizers.
Marcos Lopez de Prado (2016) · 9 citations
The paper introduces Nested Clustered Optimization (NCO) and a Monte Carlo Optimization Selection (MCOS) procedure to make mean-variance (and related) portfolio optimization far more stable out-of-sample by addressing both noisy inputs and correlation-driven “signal structure” that amplifies estimation error.
Contrarian Factor Timing Is Deceptively Difficult
Clifford Asness et al. (2017) — Journal of Portfolio Management · 42 citations
Tactical timing of investment factors (Value, Momentum, Defensive) based on their valuation spreads fails to outperform a simple diversified strategic allocation.
Scott M Lundberg et al. (2017) — Curran Associates, Inc.
This paper proposes SHAP, a unified framework showing that many popular explanation methods are all approximations to a single, uniquely justified set of feature attributions (Shapley values) with desirable properties like local accuracy and consistency.
Jing Lei et al. (2017) · 1113 citations
This paper introduces conformal prediction, a distribution-free method for constructing prediction sets in high-dimensional regression, offering finite-sample validity without strong assumptions on the data distribution or the regression estimator.
Advances in Financial Machine Learning
Marcos Lopez de Prado (2018) — John Wiley & Sons · 106 citations
Zachariah Peterson (2018) · 5 citations
The paper shows how to plug a Kelly (log-utility) return objective into a standard risk–return portfolio optimizer via a computationally convenient “decoupled” formulation, and solves the resulting nonlinear problem with differential evolution on a 10-stock example.
Takuya Akiba et al. (2019) · 9041 citations
This paper introduces Optuna, a hyperparameter optimization framework featuring a define-by-run API, efficient pruning, and versatile architecture, designed to address limitations in existing frameworks.
Yaniv Romano et al. (2019) · 848 citations
This paper introduces conformalized quantile regression (CQR), a method that combines conformal prediction with quantile regression to produce prediction intervals with finite sample coverage guarantees and adaptation to heteroscedasticity.
Nick Baltas and Robert Kosowski (2019) · 7 citations
Optimizes Time-Series Momentum (TSMOM) by reducing turnover ~36% via efficient volatility/trend estimation and improving post-crisis performance through dynamic correlation-based leverage.
Marcos Lopez de Prado (2019) · 4 citations
The paper proposes the “Theory-Implied Correlation” (TIC) algorithm, which fits a forward-looking correlation matrix constrained by an economically motivated hierarchy (e.g., GICS) rather than relying purely on unstable historical correlations.
Shihao Gu et al. (2020) — The Review of Financial Studies · 1802 citations
This paper benchmarks major machine-learning methods for predicting U.S. equity risk premia and shows that nonlinear models (especially shallow neural nets and trees) materially improve out-of-sample forecast accuracy and investable Sharpe ratios versus classic linear/panel regressions.
I. Elizabeth Kumar et al. (2020) · 453 citations
This paper argues that Shapley-value-based explanations for feature importance have mathematical and human-centric issues, making them unsuitable as a general solution.
Zihao Zhang et al. (2020) — The Journal of Financial Data Science · 135 citations
The paper proposes an end-to-end deep learning approach that outputs portfolio weights directly by maximizing the portfolio Sharpe ratio, and shows strong out-of-sample performance on a 4-ETF multi-asset portfolio from 2011 to April 2020.
Andrea Carta and Claudio Conversano (2020) — Frontiers in Applied Mathematics and Statistics · 5 citations
The paper shows how to implement Kelly-optimal (log-utility / growth-optimal) investing for equities and finds that constrained, periodically re-estimated Kelly portfolios can outperform mean-variance (tangent) portfolios—especially with a 2-year rolling estimation window—at the cost of higher risk and poorer drawdowns.
Ryan J. Tibshirani et al. (2020) · 617 citations
This paper extends conformal prediction to settings where training and test data distributions differ (covariate shift) by introducing a weighted approach that maintains valid prediction intervals, even when weights are estimated.
A. Sinem Uysal and John M. Mulvey (2021) — The Journal of Financial Data Science · 20 citations
The paper uses supervised ML (especially random forests) to predict recessions and equity “crash” regimes from macro data and then uses these probabilities to improve risk parity portfolios via regime-aware covariance estimation and overlay trades.
Isaac Gibbs and Emmanuel Candès (2021) · 373 citations
This paper introduces Adaptive Conformal Inference (ACI), a method for constructing prediction sets that maintain target coverage even when the underlying data distribution changes over time, which is crucial for real-world applications like finance where market behavior is non-stationary.
Kjersti Aas et al. (2021) — Artificial Intelligence · 849 citations
This paper extends the Kernel SHAP method to handle dependent features, providing more accurate approximations to Shapley values for explaining individual predictions in machine learning models.
Gautier Marti et al. (2021) — Springer International Publishing · 137 citations
A comprehensive review of correlation-based clustering and network methods (MST, PMFG, hierarchical clustering, RMT, etc.) for understanding market structure, building portfolios, and monitoring systemic risk—along with why these methods are often unstable and hard to operationalize.
Anastasios N. Angelopoulos and Stephen Bates (2022) · 944 citations
This tutorial explains how conformal prediction turns any black-box model’s heuristic uncertainty into prediction sets/intervals with finite-sample, distribution-free coverage guarantees (e.g., 90% coverage using a small calibration set).
Michael Pinelis and David Ruppert (2022) — The Journal of Finance and Data Science · 44 citations
The paper shows that a utility-maximizing market-timing rule that combines machine-learned forecasts of next-month equity premium and volatility ("reward-risk timing") materially improves Sharpe ratio, drawdowns, and investor utility versus buy-and-hold.
Portfolio Transformer for Attention-Based Asset Allocation
Damian Kisiel et al. (2023) — Springer International Publishing · 10 citations
The paper proposes a Transformer-based, end-to-end portfolio model that outputs asset weights by directly maximizing cost-adjusted Sharpe ratio (instead of forecasting returns then optimizing), and reports better risk-adjusted performance than LSTM and classical baselines across ETFs, commodities, and US stocks—especially around COVID-19.
Linear Model Selection and Regularization
Gareth James et al. (2023) — Springer
This chapter discusses methods to improve linear models using subset selection, shrinkage (ridge regression and lasso), and dimension reduction techniques to enhance prediction accuracy and model interpretability.
Gueorgui S. Konstantinov et al. (2023) — The Journal of Portfolio Management · 1 citations
This paper is a non-technical guide to using network (graph) representations of asset relationships to improve portfolio diversification decisions and systemic-risk awareness beyond what correlations and regressions show.
Jang Ho Kim et al. (2023) — The Journal of Portfolio Management · 7 citations
This paper is an educational review of why classical mean–variance portfolio optimization is fragile to estimation error and surveys practical robustness tools—shrinkage, constraints/regularization, robust optimization, distributionally robust optimization, scenario-based stochastic programming, and ML-based approaches.
Rina Foygel Barber et al. (2023) — The Annals of Statistics · 385 citations
This paper extends conformal prediction to provide valid prediction intervals even when data is not exchangeable (e.g., due to distribution drift or non-symmetric algorithms), using weighted quantiles and a novel randomization technique.
Alejandro Rodriguez Dominguez (2023) — Machine Learning with Applications · 5 citations
The paper proposes a portfolio construction method that diversifies using neural-network-estimated sensitivities of each asset’s return dynamics to a set of “common drivers,” then applies hierarchical clustering (HRP-style) on a sensitivity-distance matrix to produce weights.
Isaac Gibbs and Emmanuel Candès (2023) · 145 citations
This paper introduces a new method called Dynamically-Tuned Adaptive Conformal Inference (DtACI) for creating prediction sets that adapt to changing data distributions over time, improving upon existing methods by dynamically adjusting the step-size parameter based on historical performance.
Yongjae Lee et al. (2024) — The Journal of Portfolio Management · 14 citations
This educational overview explains how machine learning can improve both the inputs to portfolio optimization (returns, risk, similarity) and the optimization step itself, and how newer “decision-focused” and end-to-end methods move beyond the classic predict-then-optimize pipeline.
Sophia Sun and Rose Yu (2024) · 30 citations
This paper introduces CopulaCPTS, a novel conformal prediction algorithm for multi-step time series forecasting that uses copulas to model the joint uncertainty across future time steps, resulting in more calibrated and efficient confidence intervals with finite-sample validity guarantees.
James O'Donovan and Gloria Yang Yu (2024)
Jang Ho Kim et al. (2024) — The Journal of Portfolio Management · 3 citations
A survey of portfolio-management optimization models—from Markowitz mean–variance to robust, downside-risk, multiperiod, ESG, and ML-integrated formulations—explaining how objectives/constraints are written and solved in practice.
Alexandre Antonov et al. (2024) · 2 citations
The paper provides analytical (not just simulation) evidence that Hierarchical Risk Parity (HRP) produces less noisy portfolio weights and lower out-of-sample risk than Markowitz minimum-variance when the covariance matrix is estimated from finite samples.
David Buckle and Marielle de Jong (2024) — The Journal of Portfolio Management
The paper shows how Monte Carlo simulation can be used (even in Excel) to estimate portfolio risk measures and optimize portfolios under non-normal return distributions using copulas and expected-utility optimization.
Ross French (2024) — The Journal of Portfolio Management
The paper shows that the usual “dollar-neutral” (100/100) sizing of long–short equity factor portfolios is typically not Sharpe-optimal and proposes two simple scaling rules (Sharpe-max and volatility-matched) that improve both absolute and risk-adjusted performance, especially for stock-based short legs.
R. Douglas Martin et al. (2024) — The Journal of Portfolio Management
The paper shows—using fully reproducible R code/data—that minimum expected shortfall (MES) and minimum coherent second-moment (MCSM) portfolios can outperform classic minimum-variance portfolios in fat-tailed equity universes, but only if turnover is controlled.
Mark Kritzman (2024) — The Journal of Portfolio Management
The paper compares 1/N, mean–variance, and full-scale optimization for asset allocation, arguing that 1/N is an unjustifiable shortcut, mean–variance is a workable approximation, and full-scale optimization is the most flexible way to maximize investor utility but is computationally hard.
Joseph Simonian (2024) — The Journal of Portfolio Management · 1 citations
A methodological framework proposing 'Constructive Empiricism' to integrate econometrics (for bias reduction and explanation) and machine learning (for variance reduction and prediction) within a single investment process.
Pedro Castro et al. (2025) — The Journal of Portfolio Management
Static 0%/100% FX hedging leaves material performance on the table; simple dynamic hedge rules using carry, rolling FX–equity covariances, and (optionally) trend and PPP value improve long-run risk-adjusted returns and behave well in crises and inflationary regimes.
Christian Mueller-Glissmann (2025) — The Journal of Portfolio Management
The traditional 60/40 portfolio is failing due to positive equity-bond correlations driven by inflation; practitioners must shift to dynamic asset allocation using machine learning and broader diversification into real assets and FX overlays.
Nicholas McLoughlin (2025) — The Journal of Portfolio Management
The paper explains why “stay domestic” and “fully hedge everything” are often inefficient responses to currency risk in global portfolios, and shows how currency risk premia (carry/value/momentum) can be used in overlays or dynamic hedging—subject to sharp performance decay under real-world constraints.
The Elements of Quantitative Investing
Giuseppe A. Paleologo (2025) — John Wiley & Sons
Sophia Sun and Rose Yu (2025) · 1 citations
This paper introduces Conformal Prediction for Time-series with Change points (CPTC), a novel algorithm that integrates a model to predict underlying states with online conformal prediction to provide uncertainty quantification for time series data with change points, demonstrating improved validity and adaptivity compared to state-of-the-art baselines.
Yizhan Shu and John M. Mulvey (2025) — The Journal of Portfolio Management · 2 citations
A dynamic factor allocation strategy using Sparse Jump Models (SJM) to identify active return regimes improves the Information Ratio from 0.05 to ~0.45 compared to an equal-weighted benchmark.
Xuefeng Gao et al. (2025) · 1 citations
The paper proposes a factor-conditioned diffusion model that generates next-day cross-sectional return distributions for many stocks and uses those samples to drive daily mean–variance portfolio optimization, improving performance on China A-shares versus standard moment estimators.
Jamil Baz et al. (2025) — The Journal of Portfolio Management · 1 citations
A framework for estimating asset sensitivity to unexpected inflation and growth shocks, demonstrating that while equities and bonds suffer from inflation shocks, commodities and specific portfolio constraints can mitigate these risks.
Marcos López de Prado et al. (2025) — The Journal of Portfolio Management · 1 citations
The paper argues that you cannot compute a truly efficient portfolio frontier with a purely correlational (associational) factor model—efficient portfolio construction requires a causally specified factor model, otherwise optimization can systematically produce the wrong trades.
Vincent Tan and Stefan Zohren (2025) — The Journal of Portfolio Management · 8 citations
The paper proposes a scalable way to estimate large, exponentially-weighted covariance matrices by cross-validating eigenvalues to reduce overfitting, improving out-of-sample portfolio risk/IR in large equity universes.
Machine Learning for Probabilistic Prediction (PhD thesis, VALERY MANOKHIN)
Valery Manokhin (2022)
This thesis introduces novel methods for producing well-calibrated probabilistic predictions for machine learning classification and regression problems, addressing the issue of overconfident predictions from modern algorithms.
Søren Johansen (1991) — Econometrica · 11951 citations
This paper derives maximum-likelihood estimation and likelihood-ratio tests for cointegration rank and cointegration-vector restrictions in Gaussian VARs, yielding the “Johansen tests” based on a canonical-correlation eigenvalue problem with nonstandard asymptotic critical values.
Eugene F. Fama and Kenneth R. French (1993) — Journal of Financial Economics · 25101 citations
Fama and French propose and test a five-factor model—three equity factors (market, size, value) and two bond factors (term, default)—and show it explains most comovement in stock and bond returns and largely explains average returns across diversified portfolios.
R. N. Mantegna (1999) — Computer Physics Communications · 47 citations
The paper shows how to extract economically meaningful “sector-like” structure from stock price time series by converting correlations into distances and building a minimum spanning tree (MST) / ultrametric hierarchy.
R. Tyrrell Rockafellar and Stanislav Uryasev (2000) — The Journal of Risk · 6631 citations
The paper shows how to minimize Conditional Value-at-Risk (CVaR) by solving a convex optimization problem (often a linear program under scenarios), while simultaneously producing the corresponding Value-at-Risk (VaR).
Robert Almgren and Neil Chriss (2001) — The Journal of Risk · 564 citations
The paper derives closed-form optimal trade schedules that minimize a chosen tradeoff between market-impact costs and volatility risk, producing an “efficient frontier” of execution strategies and introducing liquidity-adjusted VaR (L‑VaR).
Ananth Madhavan (2002) — Financial Analysts Journal · 47 citations
A practitioner-focused synthesis of market microstructure showing how order flow, trading protocols, and transparency shape prices, liquidity, and transaction costs—and why these frictions matter for execution, factor models, and corporate/international finance.
André F. Perold (2004) — Journal of Economic Perspectives · 116 citations
This paper explains the intuition, derivation, and practical uses of the CAPM—showing that only non-diversifiable (market) risk should command a risk premium and that beta links expected returns to market risk.
Eugene F. Fama and Kenneth R. French (2004) — Journal of Economic Perspectives · 2104 citations
Fama and French review CAPM theory and decades of evidence, concluding CAPM fails empirically badly enough that common real-world uses (cost of capital, performance evaluation) are largely invalid.
Extending the Fundamental Law of Investment Management
Danielle Trichilo and Jeffrey L. Braun (2005) · 7 citations
The paper extends the Fundamental Law of Active Management by showing how constraints, turnover, and transaction costs affect the Information Ratio—and why allowing some shorting can materially improve implementation efficiency (transfer coefficient).
Factor Models of Asset Returns
Gregory Connor and Robert Korajczyk (2009) · 17 citations
This paper surveys how factor models decompose asset returns into pervasive and idiosyncratic components, and explains major model classes (statistical, macro, characteristics), estimation methods, and how to choose the number of factors.
Victor DeMiguel et al. (2009) — The Review of Financial Studies · 3089 citations
Across seven datasets and 14 portfolio-optimization variants, the paper finds that equal-weighting (1/N) is very hard to beat out of sample because estimation error in means/covariances overwhelms theoretical gains from mean-variance optimization.
Robert Bianchi et al. (2009) — University of Queensland · 9 citations
A Gatev-style pairs-trading strategy applied to 27 commodity futures (1990–2008) produces statistically significant, market-neutral excess returns—especially in Energy—supporting the idea that arbitrageurs are paid for enforcing the Law of One Price.
Marco Avellaneda and Jeong-Hyun Lee (2010) — Quantitative Finance · 325 citations
The paper builds and backtests market-neutral US equity mean-reversion (“stat arb”) strategies where signals come from stock residuals after removing systematic risk via PCA factors or sector ETFs, showing Sharpe ratios around 1–1.5 pre-cost with notable degradation after 2002 and a severe drawdown during August 2007.
Trading in the presence of cointegration
Alexander Galenko et al. (2012) — The Journal of Alternative Investments · 54 citations
The paper derives and backtests a dollar-neutral long/short trading strategy that exploits mean reversion implied by cointegration, proving positive expected profit under idealized conditions and showing stronger empirical performance with weekly (vs daily) sampling for a 4-ETF global equity set.
David H. Bailey and Marcos Lopez de Prado (2012) · 147 citations
The paper replaces “point-estimate Sharpe ratio” with a probability-of-skill metric (Probabilistic Sharpe Ratio, PSR) that adjusts for non-normal returns and track-record length, and extends this idea to a portfolio-selection concept called the Sharpe Ratio Efficient Frontier (SEF).
João Caldeira and Guilherme V. Moura (2013) — SSRN Electronic Journal · 110 citations
The paper builds and backtests a market-neutral pairs trading strategy for Brazilian equities that selects stock pairs via cointegration and ranks them by in-sample Sharpe, achieving ~16% annual excess returns with low market correlation out-of-sample (2006–2012).
Seven Sins of Quantitative Investing
Yin Luo et al. (2014)
A comprehensive empirical audit of seven common backtesting biases, demonstrating how errors in data handling (survivorship, look-ahead) and modeling (outliers, signal decay) can invert strategy performance from profitable to disastrous.
Robert Novy-Marx and Mihail Velikov (2016) — The Review of Financial Studies · 513 citations
The paper shows that many well-known equity anomalies look attractive before costs but lose most (or all) of their profitability after realistic trading costs—unless you redesign the strategy using a simple buy/hold “sS-rule” that sharply reduces turnover.
Zhengyao Jiang et al. (2017) · 421 citations
The paper proposes a model-free deep reinforcement learning framework that directly outputs continuous portfolio weights (not buy/sell signals), and shows it strongly outperforms many classic online portfolio selection baselines on 30-minute cryptocurrency backtests—even after 0.25% per-trade commissions.
Alan Moreira and Tyler Muir (2017) — The Journal of Finance · 381 citations
Scaling factor exposures each month by the inverse of last month’s realized variance produces large alphas and materially higher Sharpe ratios across many factors because volatility forecasts risk much more than it forecasts expected returns.
The 'Fundamental Law of Active Management' is No Law of Anything
Richard O. Michaud et al. (2017)
The paper argues—using intuition plus Monte Carlo “proofs”—that the Grinold/Kahn Fundamental Law (IR ≈ IC·√BR) is a misleading guide for real-world optimized active management because it ignores estimation error and economically necessary constraints.
Christopher Krauss (2017) — Journal of Economic Surveys · 197 citations
A structured survey of pairs trading/statistical arbitrage research that compares five major modeling frameworks and highlights what actually survives transaction costs and other market frictions.
Marco Avellaneda (2019) · 19 citations
The paper proposes Hierarchical PCA (HPCA), a sector-aware alternative to PCA that produces interpretable risk factors while preserving essentially the same information content as PCA on S&P 500 correlations.
Zihao Zhang et al. (2019) · 245 citations
The paper trains deep reinforcement learning agents to output futures trading positions directly (rather than forecasting returns), and shows they beat classic time-series momentum baselines on 50 liquid futures from 2011–2019—even with large transaction costs.
CORRGAN: Sampling Realistic Financial Correlation Matrices Using Generative Adversarial Networks
G. Marti (2020) · 43 citations
The paper introduces CorrGAN, a GAN-based method to generate synthetic but realistic financial correlation matrices that reproduce key empirical “stylized facts,” enabling better stress testing and more falsifiable empirical comparisons.
Peter Akioyamen et al. (2020) — Proceedings of the First ACM International Conference on AI in Finance · 8 citations
The paper proposes a PCA + k-means + classifier pipeline to detect U.S. macro “regimes” (crisis vs non-crisis) and shows how regime predictions can drive profitable tail-hedging and tactical allocation strategies out-of-sample (2014–2020).
Lin Cong et al. (2020) — SSRN Electronic Journal · 32 citations
The paper proposes an RL-based, Transformer-style portfolio optimizer with cross-asset attention that targets investor objectives (e.g., Sharpe) directly and then “distills” the black-box strategy into economically interpretable linear and text-topic explanations.
Xavier Gabaix and Ralph S. J. Koijen (2021) · 181 citations
The aggregate stock market is highly inelastic, meaning a $1 capital inflow increases market valuation by approximately $5, suggesting that flows rather than fundamental news drive the majority of market volatility.
Taylan Kabbani and Ekrem Duman (2022) — IEEE Access · 85 citations
The paper builds a realistic-ish stock-trading reinforcement-learning environment (with technical indicators, FinBERT headline sentiment, and transaction costs) and shows a TD3 agent can reach a 2.68 Sharpe ratio on an out-of-sample 10-stock test period.
Aaron Brixton et al. (2022) — AQR Alternative Thinking · 8 citations
This paper demonstrates that the stock-bond correlation is driven by the relative volatility of inflation versus growth, warning that a return to high inflation uncertainty could flip the correlation positive and increase 60/40 portfolio risk by 20%.
Kevin Khang (2022) — The Journal of Portfolio Management · 1 citations
Risk-model volatility forecasts depend heavily on the estimation lookback window, and a practical way to become “regime-aware” is to monitor real-time dispersion across forecasts from different calibrations and switch between slow- and fast-moving models when misspecification risk is high.
Joseph Simonian (2022) — The Journal of Portfolio Management
The paper proposes a way to make portfolio decisions when you can’t confidently identify the true economic causal model, by selecting the “most central” causal graph using a numeric distance between candidate causal schemes and then optimizing trades to minimize that distance.
Joseph A. Cerniglia and Frank J. Fabozzi (2022) — The Journal of Portfolio Management
A practitioner-oriented overview of how modern market structure, transaction-cost measurement, and optimal execution models determine whether an investment strategy’s paper alpha survives real-world implementation.
Olaf Korn et al. (2022) — The Journal of Portfolio Management · 6 citations
The paper unifies most popular drawdown risk measures under a single “weighted drawdown” framework and shows empirically (via MSCI World portfolio simulations) that different drawdown definitions can rank strategies very differently and differ in how well they detect manager skill.
Andrew Ang et al. (2023) — The Journal of Portfolio Management
A panel of industry experts defends the persistence of the Value factor despite historic drawdowns, argues for integrating ESG as alpha signals rather than standalone factors, and outlines the role of NLP and machine learning in modern factor construction.
Junkyu Jang and NohYoon Seong (2023) — Expert Systems with Applications · 80 citations
The paper proposes a DDPG-based portfolio optimizer that explicitly fuses modern portfolio theory (a 252-day correlation/covariance view) with technical-indicator signals via a tensor + Tucker decomposition architecture, and reports better Sharpe/return/drawdown than several baselines and prior DRL approaches on Dow constituents.
Milena Vuletić and Rama Cont (2023) · 12 citations
VolGAN is a conditional GAN trained on SPX options data to simulate realistic one-day-ahead joint scenarios of the SPX return and the full implied volatility surface while controlling static arbitrage via scenario reweighting, and it can be used to build data-driven option hedges that outperform delta and delta–vega hedging in their case study.
Michele Aghassi et al. (2023) — The Journal of Portfolio Management · 6 citations
A comprehensive defense of factor investing (Value, Momentum, Carry, Defensive) arguing that recent underperformance is consistent with historical risk profiles rather than structural decay or overcrowding.
Rajeev Bhargava et al. (2023) — The Journal of Portfolio Management · 4 citations
The paper builds daily, media-based “narrative” indicators (coverage intensity + sentiment) and shows they explain and sometimes predict market moves, and can be used to improve asset allocation and to build portfolios with explicit exposure to a narrative (e.g., COVID-19 recovery).
Koray D. Simsek (2023) — The Journal of Portfolio Management · 3 citations
The paper is a practitioner-oriented overview of Monte Carlo simulation—how to generate scenarios, model uncertainty and correlation, and summarize outputs—with concrete finance examples in option pricing, portfolio insurance, and VaR/ES risk measurement.
Tom Liu and Stefan Zohren (2023) · 3 citations
The paper proposes Multi-Factor Inception Networks (MFIN), an end-to-end deep learning framework that ingests many crypto “factors” (price, volume, on-chain, tweets, Google Trends) and directly outputs portfolio weights to maximize Sharpe, producing returns that are largely uncorrelated to classic momentum/reversion rules and holding up better in 2022–2023.
Peng Liu (2023) — The Journal of Portfolio Management · 1 citations
The paper shows how Bayesian optimization (BO) can tune trading-strategy parameters (treated as a noisy black-box Sharpe-ratio function) more efficiently than manual or random search, using pairs trading as a worked example.
Giuseppe Paleologo (2023) — The Journal of Portfolio Management · 1 citations
The paper gives an exact, implementable decomposition of a strategy’s idiosyncratic information ratio into (selection skill × effective breadth) plus a sizing skill term, using Herfindahl-based breadth rather than “√n”.
David Blitz (2023) — The Journal of Portfolio Management · 1 citations
A strategic review of factor investing that argues for specific modernizations—including intangible-adjusted value, continuous trading, and short-term signals—to overcome the performance stagnation of generic Fama-French factors.
R. Douglas Martin et al. (2023) — The Journal of Portfolio Management · 1 citations
The paper shows how modern robust regression and robust covariance estimators can materially change factor-model inference and portfolio risk estimates by down-weighting a small fraction of influential outliers that can severely distort least-squares and sample-covariance results.
Bradford Cornell (2023) — The Journal of Portfolio Management · 1 citations
This paper explains Bayes’ rule and argues that “Bayesian thinking” (explicitly weighing priors against noisy evidence) is a practical lens for portfolio choice, manager evaluation, and especially for interpreting factor/anomaly research under data-mining risk.
Eric Sorensen et al. (2023) — The Journal of Portfolio Management · 1 citations
Tracking error (TE) constraints can be counterproductive because a cap-weighted benchmark like the S&P 500 changes its diversification (concentration) over time, implying skilled managers should vary TE rather than keep it fixed.
Stephen Marra (2023) — The Journal of Portfolio Management
A practitioner-focused survey comparing common volatility forecasting models (historical, ARMA/GARCH, and option-implied) and showing why relatively simple, well-designed historical models can be robust inputs for volatility targeting and risk-parity allocation.
Ding Liu (2023) — The Journal of Portfolio Management
The paper proposes a common, implementation-aware framework to compare equity downside-protection approaches (CPPI, vol targeting, and option overlays) and finds they deliver essentially the same long-run downside-vs-upside trade-off from 1996–2020, but differ sharply by drawdown/recovery stage depending on whether their payoff is concave or convex.
Sid Browne et al. (2023) — The Journal of Portfolio Management
The paper proposes statistically testable definitions of “timing” (win frequency) and “sizing” (win/loss magnitude) skill for systematic strategies, and shows which strategy types look better or worse in inflation and recession regimes.
Dirk G. Baur and Thomas Dimpfl (2023) — The Journal of Portfolio Management
Using randomized portfolios of large US stocks, the paper tests the famous rule “cut your losses and let your profits run” and finds that threshold-based “cut” strategies consistently underperform simple buy-and-hold—even during crises—because losers often rebound while winners exhibit stronger persistence.
Ashwin Alankar et al. (2023) — The Journal of Portfolio Management · 1 citations
Using 1871–2022 US monthly data, the paper identifies 23 equity drawdowns ≥15%, clusters them into 5 pre-crisis “environments,” and shows a simple 2-variable model can usually tell whether the next drawdown will be a fast “gamma” crash or a slow “delta” grind—guiding tail-hedge choice.
Benton Chambers et al. (2024) — The Journal of Portfolio Management
This paper demonstrates that simple, model-free definitions of Value, Carry, Quality, and Momentum generate robust risk-adjusted returns across US Investment Grade, High Yield, and Emerging Markets, with a multi-factor US High Yield portfolio achieving an Information Ratio of 0.68 after transaction costs.
Yangyang Yu et al. (2024) · 112 citations
FINCON is a GPT-4-based manager–analyst multi-agent trading system that uses CVaR-based real-time risk alerts plus “conceptual verbal reinforcement” across training episodes to improve sequential trading decisions for both single stocks and small portfolios.
Gian Luca Tassinari et al. (2024) — The Journal of Portfolio Management · 1 citations
A practical tutorial on defining and measuring market risk (volatility, VaR, ES, and expectile-based VaR) and using these measures to build minimum-risk portfolios, illustrated with a 4-asset bond/equity case study.
Alejandro Rodriguez Dominguez (2024) · 1 citations
The paper derives a “conditional risk-neutral” PDE system for a portfolio by conditioning constituent returns on a small set of common causal drivers (modelled via Gaussian copulas), and proposes monitoring PDE deviations as a dynamic risk-management signal.
Gueorgui S. Konstantinov (2025) — The Journal of Portfolio Management
A framework for transforming currency from a passive hedging byproduct into an active alpha source by integrating carry, value, trend, and volatility styles via risk budgeting.
Carter Davis et al. (2025) — The Journal of Portfolio Management
The paper argues that “style factors” (value/size/momentum sorts) leave substantial Sharpe ratio and alpha unexplained, and proposes “sectional factors”—demand/flow-driven factor portfolios built from many characteristics and machine learning—to materially expand the mean–variance efficient frontier out of sample.
Frank J. Fabozzi (2025) — The Journal of Portfolio Management
López de Prado argues that finance requires specialized 'Alpha Assembly Lines' and distinct ML algorithms (like HRP and DSR) to mine 'microscopic alpha' and avoid the overfitting traps of standard econometrics and backtesting.
Marcos López de Prado and Frank J. Fabozzi (2025) — The Journal of Portfolio Management
A comprehensive survey detailing ten specific areas where machine learning outperforms traditional econometrics in finance, ranging from hierarchical portfolio construction to meta-labeling for bet sizing.
Rashmi Malhotra and D. K. Malhotra (2025) — The Journal of Portfolio Management
The paper shows how unsupervised clustering (k-means, DBSCAN, hierarchical) can group stocks by realized risk–return behavior (return, volatility, Sharpe) rather than sector labels, enabling behavior-aware diversification, defensive positioning, and thematic idea generation.
Ricky Cooper and Zixuan Jiao (2025) — The Journal of Portfolio Management
A momentum-neutralized, linear combination of Value, Profitability, and Momentum factors outperforms sequential sorting strategies, delivering higher risk-adjusted returns with significantly lower turnover.
Adil Rengim Cetingoz and Charles-Albert Lehalle (2025)
The paper argues that synthetic financial return data from generic generative models can mislead portfolio/risk conclusions—especially for long-short portfolios—because (i) you cannot “create information” beyond the original sample size and (ii) standard generative losses learn the wrong directions (high-variance PCs) for portfolio optimization.
Gueorgui S. Konstantinov and Frank J. Fabozzi (2025) — The Journal of Portfolio Management · 1 citations
The paper proposes CAFNITE, a causal-network framework that measures how shocks to one factor/asset propagate through a global multi-asset factor network (2001–2024), showing diversification depends on time-varying causal linkages rather than static correlations and offering early-warning diagnostics before volatility spikes.
Nusret Cakici et al. (2025) — The Journal of Portfolio Management
The paper shows that factor (anomaly portfolio) returns are meaningfully predictable in the cross section using ML on factor-level characteristics, but most of the “ML alpha” is essentially factor momentum.
Petter N. Kolm and Gordon Ritter (2025) — The Journal of Portfolio Management
An equation-free, practitioner-oriented overview of how reinforcement learning (RL) can be used to learn adaptive trading, hedging, and portfolio policies under frictions where today’s actions change tomorrow’s opportunities.
Unknown
This blog post explores how Bayesian statistics, particularly with the PyMC library, can enhance financial analysis by quantifying uncertainty, overcoming restrictive assumptions of traditional econometric models, and managing complex model structures.
Systematic Fixed- Income Investing Comes of Age
Scott DiMaggio and Bernd Wuebben
A framework for applying systematic multifactor investing to fixed income markets, emphasizing the necessity of proprietary liquidity modeling and specific factor definitions (Value, Carry, Quality) distinct from equities.
Deep Reinforcement Learning for Optimal Portfolio Allocation: A Comparative Study with Mean-Variance Optimization
Srijan Sood et al. · 16 citations
The paper runs a like-for-like backtest comparing a PPO-based deep RL allocator to a carefully-implemented mean-variance optimizer on US sector indices (2012–2021) and finds DRL delivers materially higher Sharpe, higher returns, and more stable/less-turnover allocations under the same long-only constraints.