Causal Machine Learning
Learning Objectives
- Define a causal research question in terms of treatment, outcome, estimand, and counterfactual, and use DAGs to
- Apply validation and refutation tools, including placebo tests, sensitivity analysis, and subset-stability checks, to
- Use Double Machine Learning (DML) to estimate causal effects of continuous treatments in the presence of
- Use Bayesian Structural Time-Series (BSTS) to estimate the impact of discrete events by constructing data-driven
- Use causal discovery methods such as PCMCI, NOTEARS, and VAR-LiNGAM to generate candidate structures and interpret
- Distinguish predictive signal from causal effect, and interpret cross-dataset evidence with attention to confounding
From theory to estimation
Identification and validation
Validation and refutation
Isolating factor effects with DML
Measuring event impact with Bayesian structural time-series
Causal discovery from observational data
Case study causal evidence
Summary
Related Case Studies
See where these chapter concepts get applied in end-to-end trading workflows.
ETF Cross-Asset Exposures
All six model families compared across 100 ETFs spanning 9 asset classes
Crypto Perpetuals Funding
Alternative data and non-standard frequencies in 24/7 crypto markets
NASDAQ-100 Microstructure
Intraday microstructure signals across 114 stocks at 15-minute frequency
S&P 500 Equity + Option Analytics
Combining options-derived features with equity data for multi-source prediction
US Firm Characteristics
Classic factor investing with ML on monthly fundamental data
FX Spot Pairs
Momentum and carry factors in the world's most liquid market
CME Futures
Carry signals across 30 products — data quality as the critical variable
S&P 500 Options (Straddles)
Direct options trading and why equity-style cost models fail for options
US Equities Panel
Large-scale cross-sectional prediction across 3,200 stocks with 16 walk-forward folds
01 Library Overview
02 Dowhy Causal Graph
03 Econml Dml
04 Dml Crypto Regime
05 Momentum Causal Trading
06 Fed Announcement Bsts
07 Tigramite Time Series
08 Neural Causal Discovery
09 Adia Causal Benchmark
10 Case Study Insights
11 Factor Zoo Validation
8 primer topics providing foundational concepts for this chapter.
Backdoor Adjustment and Control Selection in Causal DAGs
Before you estimate a treatment effect, you need to know which variables identify it and which ones poison it.
Bayesian Inference and MCMC for Time Series
A Bayesian time-series model produces a posterior distribution, not just a fitted line, which is why posterior uncertainty can itself become a feature.
Causality, Confounding, and Why Good Signals Can Be Misleading
A predictive relationship can be real in the data and still fail as an explanation of what would happen under intervention.
Interference, Spillovers, and SUTVA Violations in Financial Markets
In markets, one unit's treatment rarely stays politely confined to that unit.
Potential Outcomes and the Rubin Causal Model
Causal inference starts with a missing-data problem: for each unit, the outcome you most want to compare is the one you never get to observe.
Sensitivity Analysis, Placebos, and Negative Controls in Observational Finance
A causal estimate becomes usable only after you ask how it could be wrong.
What the BSTS Counterfactual Actually Learns
A BSTS event study builds a synthetic no-intervention world from two ingredients: the target series' own dynamics and its pre-event co-movement with controls. Understanding exactly what each component contributes -- and what breaks each one -- is essential for credible causal claims.
Why Orthogonalization Works: The Mechanism Behind Double Machine Learning
DML does not succeed because it runs two models instead of one. It succeeds because the Neyman-orthogonal score is locally insensitive to first-order errors in the nuisance estimates, and cross-fitting breaks the dependence between estimation error and score evaluation.
63 references cited in this chapter.
Peter Spirtes et al. (1993) — Springer · 2780 citations
Guido W. Imbens and Joshua D. Angrist (1994) — Econometrica · 5082 citations
This paper introduces the concept of Local Average Treatment Effect (LATE) as an alternative to Average Treatment Effect (ATE) when evaluating social programs, especially in cases where full random assignment is not feasible, providing a method to identify treatment effects for subpopulations influenced by an instrumental variable.
Aapo Hyvärinen et al. (2010) — Journal of Machine Learning Research · 440 citations
This paper introduces a novel method for estimating Structural Vector Autoregression (SVAR) models by leveraging the non-Gaussianity of disturbance variables, overcoming identifiability issues common in traditional SVAR estimation.
Tim Loughran and Bill Mcdonald (2011) — The Journal of Finance · 711 citations
Generic sentiment dictionaries (e.g., Harvard-IV negative words) badly misclassify tone in 10-Ks, so the paper builds finance-specific dictionaries and shows they better link 10-K language to market reactions and firm outcomes.
Alexandre Belloni et al. (2012)
Paul C. Tetlock (2014) — Annual Review of Financial Economics · 87 citations
A comprehensive review of how media content, linguistic tone, and investor attention influence asset prices, establishing the foundation for Natural Language Processing (NLP) in finance.
David H. Bailey and Marcos Lopez de Prado (2014) · 110 citations
The paper introduces the Deflated Sharpe Ratio (DSR), a statistic that adjusts performance metrics for the probability of backtest overfitting caused by multiple testing and non-normal returns.
Kay H. Brodersen et al. (2015) — The Annals of Applied Statistics · 854 citations
The paper introduces a fully Bayesian state-space (structural time-series) approach to estimate the causal effect of an intervention by forecasting a counterfactual time series (synthetic control) and comparing it to observed outcomes, implemented in the CausalImpact R package.
Alya Al Nasseri et al. (2015) — Expert Systems with Applications · 54 citations
This paper applies Decision Trees to StockTwits data to identify specific semantic terms (and term combinations) whose volume fluctuations predict DJIA price movements, finding that 'sell' terms are more predictive than 'buy' terms.
Natural Language Processing - Part I: Primer
Frank Zhao (2017)
This primer explains core NLP concepts and shows (with code) how to turn earnings-call transcripts into investable signals such as negativity-based sentiment changes and language-complexity measures linked to future returns.
Victor Chernozhukov et al. (2018) — The Econometrics Journal · 3298 citations
This paper shows how to get valid (root-N, asymptotically normal) inference for causal/structural parameters when nuisance components are learned with high-dimensional ML, by combining Neyman-orthogonal scores with cross-fitting.
Xun Zheng et al. (2018) · 1250 citations
This paper introduces a novel method called NOTEARS for learning directed acyclic graphs (DAGs) by reformulating the combinatorial optimization problem into a continuous one, enabling the use of standard numerical algorithms.
Natural Language Processing - Part II: Stock Selection
Frank Zhao and Quantamental Research (2018)
Using earnings-call transcripts, simple NLP features (sentiment and “transparency” proxies like language simplicity and numeric intensity) predict cross-sectional U.S. stock returns, with long-short spreads of ~2–4% per year that remain significant after controlling for common factors.
Yinhan Liu et al. (2019) — arXiv:1907.11692 [cs] · 28968 citations
RoBERTa shows that many post-BERT gains come from a better training recipe (more data, longer training, larger batches, no NSP, dynamic masking) rather than fundamentally new pretraining objectives, yielding state-of-the-art results on GLUE, SQuAD, and RACE.
Jakob Runge et al. (2019) — Nature Communications · 904 citations
This Perspective surveys modern methods for causal discovery from time series (beyond correlation/forecasting) and maps them to recurring real-world problems—arguing for benchmarks and clearer assumptions to make causal claims credible in complex dynamical systems.
Jacob Devlin et al. (2019) — Association for Computational Linguistics · 112230 citations
BERT introduces a Transformer encoder pre-trained with masked-token prediction (and next-sentence prediction) to learn deep bidirectional language representations that can be fine-tuned with minimal task-specific changes to achieve state-of-the-art NLP results.
Nils Reimers and Iryna Gurevych (2019) — arXiv:1908.10084 [cs] · 16856 citations
SBERT modifies BERT into a siamese/triplet architecture to produce cosine-comparable sentence embeddings that make semantic search and clustering practical (seconds instead of hours) while retaining strong accuracy on STS-style tasks.
Tim Loughran and Bill McDonald (2020) — Annual Review of Financial Economics · 169 citations
A practitioner-oriented review of how finance uses text (social media, politics, fraud) that argues “readability” metrics like the Fog Index are mis-specified for 10-Ks and should be replaced by text-based measures of firm complexity.
Guanhao Feng et al. (2020) — The Journal of Finance · 490 citations
The authors propose a Double-Selection LASSO methodology to evaluate new asset pricing factors against hundreds of existing ones, finding that while most new factors are redundant, profitability and investment factors provide significant marginal value.
Lingyun Zhao et al. (2020) — arXiv:2001.05326 [cs] · 80 citations
The paper fine-tunes RoBERTa to (1) classify sentiment in online financial text with emphasis on negative sentiment and (2) extract the key entity/entities driving that negative news using sentence matching and MRC-style QA instead of standard NER.
Alexander Reisach et al. (2021) — Curran Associates, Inc. · 182 citations
This paper demonstrates that the performance of continuous causal structure learning algorithms on simulated data can be attributed to a property called 'varsortability,' where marginal variance increases along the causal order, and that this advantage disappears when data is standardized.
Marcos Lopez de Prado (2022) · 13 citations
The paper argues that current factor investing relies on spurious associational correlations rather than causal mechanisms, and proposes adopting 'Causal Factor Investing' using causal graphs and do-calculus to distinguish true drivers of returns from statistical artifacts.
Jules H. van Binsbergen et al. (2022) · 3 citations
The paper shows that markets underreact to fraud-related cash-flow news in short-seller reports: prices drop about −4.9% on release day but continue drifting down to roughly −15% over the next year, especially when the report’s text emphasizes fraud.
Ali Shojaie and Emily B. Fox (2022) — Annual Review of Statistics and Its Application · 458 citations
This review explains what Granger causality really measures (predictive, time-ordered dependence), why naive/bivariate VAR tests can mislead, and how modern high-dimensional, nonlinear, discrete-valued, and mixed-frequency methods extend the framework.
Andrew W. Lo et al. (2023) — The Journal of Portfolio Management · 20 citations
This article surveys how NLP evolved from rule-based systems to transformer LLMs and explains what that evolution means for practical financial workflows like trading signals, risk management, and ESG/impact investing.
Martin Luk (2023) · 11 citations
A comprehensive survey detailing the technical architecture of Generative AI (Transformers, Diffusion) and its specific applications in quantitative finance, including superior sentiment analysis and synthetic data generation.
Rajeev Bhargava et al. (2023) — The Journal of Portfolio Management · 4 citations
The paper builds daily, media-based “narrative” indicators (coverage intensity + sentiment) and shows they explain and sometimes predict market moves, and can be used to improve asset allocation and to build portfolios with explicit exposure to a narrative (e.g., COVID-19 recovery).
Qianqian Xie et al. (2023) — arXiv preprint arXiv:2306.05443 · 259 citations
PIXIU open-sources an instruction-tuned finance LLM (FinMA), a 128,640-sample finance instruction dataset (FIT), and a benchmark (FLARE) spanning NLP and prediction tasks to systematically train and evaluate LLMs for finance.
Yaxuan Kong et al. (2024) — The Journal of Portfolio Management · 8 citations
This survey maps the emerging landscape of large language models (LLMs) in finance—what specialized models exist, where they can improve investment workflows, and which technical, evaluation, and ethical risks must be managed for responsible deployment.
Unknown (2024)
This blog post explains the BM25 full-text search algorithm, its components (query terms, IDF, term frequency, document length normalization), and its clever probabilistic ranking approach, concluding that BM25 scores can be compared within the same document collection.
Ali Jaffri et al. (2025) — The Journal of Portfolio Management
This article is a step-by-step, practitioner-oriented tutorial showing how FinBERT can convert firm-specific financial news headlines into an interpretable sentiment index that can be overlaid on intraday price charts to contextualize short-term market moves.
Rohan Alur et al. (2025) · 1 citations
A multi-agent LLM system achieves human superforecaster performance by combining agentic search, supervisor-based reconciliation, and statistical calibration.
Paul Glasserman et al. (2025)
Using 2.4M Reuters articles (1996–2022), the paper shows that persistent firm-level news-topic exposures and opposite intraday vs overnight price responses to the same topics explain a large share of the long-run U.S. “overnight drift” (overnight gains vs flat/negative intraday returns) and help forecast which stocks will do best overnight and worst intraday.
Emanuele Olivetti et al. (2026)
This paper analyzes the results of the ADIA Lab Causal Discovery Challenge, finding that supervised learning methods, sophisticated feature engineering, and ensemble methods significantly outperformed traditional constraint-based approaches in inferring causal relationships from synthetic data.
Man Group · 1235 citations
A roundtable of leading academics and Man Group practitioners discussing the specific challenges of applying Generative AI to quantitative finance, identifying look-ahead bias in backtesting as the primary obstacle.
BlackRock
BlackRock discusses how they've been using AI and machine learning in systematic investing for nearly two decades, focusing on LLMs for text analysis and thematic basket construction.
Tobias Preis et al. (2013) — Scientific Reports · 942 citations
This seminal paper demonstrates that increases in Google search volume for financially negative terms like 'debt' precede stock market falls, enabling a strategy that returned 326% over 7 years.
Tomas Mikolov et al. (2013) — arXiv preprint arXiv:1301.3781 · 33959 citations
The paper introduces CBOW and Skip-gram—two simple, fast neural architectures that learn high-quality word embeddings from billions of tokens and exhibit strong linear “analogy” structure (e.g., king − man + woman ≈ queen).
Ashish Vaswani et al. (2017) — arXiv:1706.03762 [cs] · 171159 citations
This paper introduces the Transformer, a sequence-to-sequence model that replaces recurrence and convolutions with multi-head self-attention, achieving state-of-the-art translation quality with much faster, highly parallel training.
Alexis Conneau and Douwe Kiela (2018) — European Language Resources Association (ELRA) · 677 citations
SentEval is a standardized, easy-to-use toolkit that evaluates universal sentence embeddings across a community-agreed suite of transfer tasks, reducing pipeline variability and making results more comparable.
Victor Sanh et al. (2020) — arXiv:1910.01108 [cs] · 9380 citations
The paper shows how to pre-train a smaller BERT-like model using knowledge distillation so it keeps almost all of BERT’s accuracy while being much smaller and faster—making on-device NLP practical.
Lin Cong et al. (2020) — SSRN Electronic Journal · 32 citations
The paper proposes an RL-based, Transformer-style portfolio optimizer with cross-asset attention that targets investor objectives (e.g., Sharpe) directly and then “distills” the black-box strategy into economically interpretable linear and text-topic explanations.
Gene Ekster and Petter N. Kolm (2020) · 5 citations
A comprehensive guide to the alternative data ecosystem, detailing a specific preprocessing pipeline (entity tagging, stabilization, debiasing) that reduces revenue prediction error from 88% to 2.6%.
Shuohang Wang et al. (2021) · 314 citations
The paper evaluates GPT-3 as a cheap “pseudo-labeler” to annotate training data for smaller deployable NLP models, showing large labeling cost savings (50%–96%) and further gains when mixing GPT-3 and human labels under a fixed budget.
Taylan Kabbani and Ekrem Duman (2022) — IEEE Access · 85 citations
The paper builds a realistic-ish stock-trading reinforcement-learning environment (with technical indicators, FinBERT headline sentiment, and transaction costs) and shows a TD3 agent can reach a 2.68 Sharpe ratio on an out-of-sample 10-stock test period.
Yongjae Lee et al. (2023) — The Journal of Portfolio Management · 18 citations
A comprehensive survey categorizing machine learning models by their specific applications in asset management, ranging from tree-based models for tabular data to reinforcement learning for hedging.
Jason Wei et al. (2023) · 16721 citations
The paper shows that giving large language models a few examples that include intermediate reasoning steps (“chain-of-thought”) can dramatically improve accuracy on multi-step reasoning tasks without finetuning.
Chang Gong et al. (2023) · 63 citations
A comprehensive survey that unifies causal discovery methods for both regularly-sampled multivariate time series and irregular event sequences, and lays out practical evaluation resources plus future directions (amortized/supervised causal discovery and causal representation learning).
Alejandro Lopez-Lira (2023) · 34 citations
The paper uses topic modeling on 10-K “Item 1A Risk Factors” text to infer firm-specific risk exposures, builds tradable risk-mimicking portfolios and a parsimonious factor model from them, and shows these “firm-identified risk factors” add pricing information beyond standard factor models.
Andrew Ang et al. (2023) — The Journal of Portfolio Management
A panel of industry experts defends the persistence of the Value factor despite historic drawdowns, argues for integrating ESG as alpha signals rather than standalone factors, and outlines the role of NLP and machine learning in modern factor construction.
T. Clifton Green and Shaojun Zhang (2024) — The Journal of Portfolio Management · 2 citations
A comprehensive survey of the alternative data landscape, categorizing sources into firm-released, government, attention, and third-party data, while analyzing the regulatory drivers and implementation challenges like alpha decay.
Qianqian Xie et al. (2024) — Advances in Neural Information Processing Systems · 123 citations
FinBen is a large open-source benchmark (42 datasets, 24 tasks) designed to measure what today’s LLMs can and cannot do in real financial NLP, forecasting, risk, and trading/agent settings.
Dhagash Mehta et al. (2025) — The Journal of Portfolio Management · 1 citations
A comprehensive tutorial on replacing rigid financial classifications (like GICS or style boxes) with adaptive, multimodal machine learning techniques to construct operationally valid peer groups.
Andrew Chin (2025) — The Journal of Portfolio Management · 1 citations
Large Language Models (LLMs) are converging investment styles by giving discretionary managers scalable 'breadth' and systematic managers qualitative 'depth', creating a hybrid 'Iron-Person' investor model.
Alejandro Lopez-Lira et al. (2025)
LLMs can appear to “forecast” economics/markets pre-cutoff because they often memorize realized historical outcomes, making pre-cutoff forecast evaluation fundamentally non-identifiable and prone to lookahead bias.
Iro Tasitsiomi and Yijie Wang (2025) — The Journal of Portfolio Management
A practitioner's framework for 'quantamental' investing that categorizes how alternative data and AI models create new factors to augment, rather than replace, fundamental analysis.
Guido Baltussen et al. (2025) — The Journal of Portfolio Management
This practitioner-oriented paper explains how asset managers can convert large-scale text (filings, earnings calls, news) into investable signals using NLP—from bag-of-words to LLM embeddings—while managing leakage, bias, hallucinations, translation, and cost.
Chanyeol Choi et al. (2025) · 16 citations
FINDER is a 5,703-example expert-annotated dataset of real, ambiguous finance search queries grounded in S&P 500 10‑K evidence, designed to realistically benchmark retrieval-augmented generation (RAG) for financial question answering.
Ming Deng et al. (2025) — The Journal of Portfolio Management · 1 citations
Four vendor news-sentiment feeds are only weakly correlated, and averaging them into a single monthly firm-level signal produces a stronger, robust cross-sectional return premium that looks like slow-moving fundamental information rather than compensation for risk.
Unknown
This article from Anthropic discusses best practices for building LLM agents, emphasizing simple, composable patterns over complex frameworks, and provides guidance on when to use workflows versus autonomous agents.
Hoogkamer
Termboard is a user-friendly, browser-based tool that simplifies the creation of knowledge graphs and semantic models for various applications, requiring no specialized expertise or complex installations.
Unknown
This is a sign-in page for alphaXiv, a platform for exploring AI research, offering features like AI-powered text explanations and summaries.
Unknown
A 2025 industry survey revealing that 67% of investment firms now use alternative data, with rapid adoption of AI (61%) and significant budget increases, though the sample is heavily skewed toward Private Equity.