MLOps and Governance
Learning Objectives
- Distinguish technical pipeline divergence from statistical performance decay and choose the corresponding diagnostic
- Build a live-monitoring framework that combines data-integrity gates, rolling performance metrics, backtest-to-live
- Apply drift diagnostics to production artifacts, including PSI, K-S, SHAP-based feature monitoring, and online
- Design a safe model-update workflow using shadow mode, champion-challenger evaluation, explicit promotion criteria,
- Implement multi-level circuit breakers across trade, strategy, portfolio, and system layers, with clear recovery and
- Evaluate and right-size the supporting MLOps stack, including feature stores, data versioning and lineage, model
Two sources of live trading failure
This section establishes the fundamental distinction between technical failures (pipeline divergence where same inputs produce different outputs, addressed by verification in Chapter 25) and statistical failures (correct implementation but decayed predictions due to overfitting, look-ahead bias, regime change, or alpha crowding). The distinction matters operationally because treating statistical decay as a bug wastes debugging time, while treating bugs as decay leads to unnecessary model changes. Four mechanisms of statistical decay are identified, with the monitoring framework in subsequent sections designed around the assumption that decay will happen and must be detected early enough to respond.
Performance monitoring
The section builds a multi-layered monitoring framework starting with data integrity gates that catch silent data defects before they propagate, then rolling metrics (Sharpe ratio, IC, hit rate, drawdown) computed across multiple trailing windows to reveal trends that aggregate statistics hide. Tiered alert thresholds (watch, warning, critical) with defined response actions prevent both under-reaction and alert fatigue, while the backtest-to-live realization ratio tracks whether live performance matches expectations over time. Execution-quality monitoring for slippage, spread costs, fill ratios, and latency is treated as essential alongside model metrics, since worsening execution can masquerade as model decay.
Drift detection
This section provides diagnostic tools that identify what changed when performance monitoring detects problems, covering three drift types: data drift measured by PSI and K-S tests on input feature distributions, feature drift tracked through SHAP value monitoring that reveals importance shifts even when distributions remain stable, and concept drift detected by ADWIN and DDM algorithms applied to prediction error streams. A four-quadrant diagnostic table crossing drift detection with performance decay guides response: drift with no decay means the model is robust, decay without detected drift means monitoring coverage is incomplete, and both together means retraining on recent data is warranted.
Safe model updates
The section presents a disciplined model update workflow combining scheduled and triggered retraining, shadow mode evaluation where challenger models process live data without trading, capital-capped A/B testing with gradual allocation increases, and explicit rollback procedures tested before deployment. Statistical rigor in promotion decisions requires the deflated Sharpe ratio or bootstrap comparison accounting for estimation error and multiple testing, with minimum effect size thresholds (0.2-0.3 Sharpe improvement) below which promotion is rejected regardless of statistical significance. When multiple challengers compete, White's Reality Check alongside multiple-testing corrections prevents selection bias from producing false promotion decisions.
Circuit breakers and safety
This section addresses sudden failure through four hierarchical circuit breaker levels: per-trade order validation, per-strategy exposure limits, portfolio-wide aggregate risk controls, and system-level infrastructure health monitoring, each operating independently so a breach at any level halts the relevant scope. Loss-based breakers with daily, weekly, and maximum drawdown limits are complemented by position concentration limits, anomaly-based triggers for extreme market conditions, and software circuit breakers using the CLOSED/OPEN/HALF_OPEN state machine pattern to prevent cascading infrastructure failures. Recovery procedures require explicit resume criteria, gradual restart at reduced capacity, and logged manual overrides with time limits and mandatory after-the-fact review.
MLOps infrastructure overview
The section surveys infrastructure tools supporting production governance, organized by maturity level: feature stores (Feast) for preventing training-serving skew through consistent feature definitions across offline training and online inference, data versioning (DVC) with run manifests for exact reproducibility, model registries (MLflow) for experiment tracking and staged deployment with auditability for regulatory review under SR 11-7. The right-sizing guidance progresses from minimal stacks suitable for solo practitioners through intermediate setups with Prometheus-Grafana monitoring to mature Kubernetes-based deployments, arguing that monitoring and safety controls from earlier sections matter more than tooling choices and should be established first.
Summary
01 Drift Monitoring
02 Online Drift Detection
03 Safe Model Rollout
04 Circuit Breakers
05 Feast Feature Store
05B Feast Live
06 Mlflow Experiments
4 primer topics providing foundational concepts for this chapter.
Champion-Challenger Evaluation and Shadow Mode in Trading Systems
"When comparing two strategies' Sharpes, prioritize dependence-aware bootstrap tests." -- Ledoit and Wolf (2008)
Drift Detection and Trigger Design
A drift monitor is useful only when you know what statistic it watches, what kind of change it can see, and how many false alarms you are willing to tolerate.
Hypothesis Testing and P-Values
How hypothesis tests turn noisy evidence into a structured decision, and how to read p-values without treating them as proof.
Training-Serving Skew, Point-in-Time Joins, and Feature Stores
"ML systems have all of the maintenance problems of traditional code plus an additional set of ML-specific issues." -- Sculley et al. (2015)
79 references cited in this chapter.
Ana L. C. Bazzan et al. (2004) — Springer Berlin Heidelberg · 1614 citations
This paper introduces a method for detecting concept drift in online learning by monitoring the error rate of a learning algorithm and adapting the model when significant increases in error suggest a change in the underlying data distribution.
Albert Bifet and Ricard Gavaldà (2007) — Society for Industrial and Applied Mathematics · 1753 citations
{Board of Governors of the Federal Reserve System (2011) · 51 citations
The foundational regulatory framework establishing the industry standard for Model Risk Management (MRM), mandating independent validation, effective challenge, and governance for all financial models.
David H. Bailey and Marcos Lopez de Prado (2014) · 110 citations
The paper introduces the Deflated Sharpe Ratio (DSR), a statistic that adjusts performance metrics for the probability of backtest overfitting caused by multiple testing and non-normal returns.
D. Sculley et al. (2015) — Curran Associates, Inc. · 1336 citations
This paper discusses the concept of technical debt in machine learning systems, highlighting the often-overlooked maintenance costs and ML-specific risk factors that can lead to long-term problems.
...and the Cross-Section of Expected Returns
Campbell R. Harvey et al. (2016) — Review of Financial Studies · 1838 citations
Due to extensive data mining in asset pricing ('the factor zoo'), the standard t-statistic threshold of 2.0 is insufficient; this paper mathematically demonstrates that a t-statistic > 3.0 is required to establish true significance.
Does Academic Research Destroy Stock Return Predictability?
R. David McLean and Jeffrey Pontiff (2016) — Journal of Finance · 851 citations
Stock return predictors published in academic journals lose 58% of their performance after publication due to a combination of statistical bias (data mining) and arbitrage trading.
Scott M Lundberg et al. (2017) — Curran Associates, Inc.
This paper proposes SHAP, a unified framework showing that many popular explanation methods are all approximations to a single, uniquely justified set of feature attributions (Shapley values) with desirable properties like local accuracy and consistency.
Advances in Financial Machine Learning
Marcos Lopez de Prado (2018) — John Wiley & Sons · 106 citations
Jie Lu et al. (2018) — IEEE Transactions on Knowledge and Data Engineering · 1684 citations
This paper reviews methodologies and techniques for concept drift detection, understanding, and adaptation in machine learning, providing a framework and discussing datasets and research directions.
Victor Sanh et al. (2020) — arXiv:1910.01108 [cs] · 9380 citations
The paper shows how to pre-train a smaller BERT-like model using knowledge distillation so it keeps almost all of BERT’s accuracy while being much smaller and faster—making on-device NLP practical.
Stefan Studer et al. (2021) — Machine Learning and Knowledge Extraction · 235 citations
A comprehensive process model extending CRISP-DM to include Quality Assurance (QA) and a specific 'Monitoring and Maintenance' phase to address the 75-85% failure rate of ML projects.
Olaf Korn et al. (2022) — The Journal of Portfolio Management · 6 citations
The paper unifies most popular drawdown risk measures under a single “weighted drawdown” framework and shows empirically (via MSCI World portfolio simulations) that different drawdown definitions can rank strategies very differently and differ in how well they detect manager skill.
Fabian Hinder et al. (2023) · 17 citations
This paper surveys concept drift detection methods in unsupervised data streams, focusing on monitoring and anomaly detection, providing a taxonomy, localization techniques, and standardized experiments.
Andrei Paleyes et al. (2023) — ACM Computing Surveys · 546 citations
This survey paper identifies and categorizes the challenges faced when deploying machine learning models into production environments across various industries, highlighting issues at each stage of the ML deployment workflow.
Agostino Capponi et al. (2025)
This paper introduces the nonstationarity-complexity tradeoff in financial return prediction, proposing an adaptive model selection method (ATOMS) that jointly optimizes model complexity and training window size to significantly improve out-of-sample R^2 and trading returns, especially during economic recessions.
Samir Varma (2025) — The Journal of Portfolio Management
Fixed drawdown-triggered de-risking (e.g., “cut risk at −10%”) often worsens outcomes by forcing exits before recoveries, and a context-aware framework built around a coherent drawdown-adjusted metric (CDAP) better identifies true crisis risk.
Marcos Lopez de Prado et al. (2025)
This paper is a practical and statistical “user manual” for Sharpe ratio inference, showing how to correct for non-Normality, serial correlation, small samples, and multiple testing so you don’t mistake noise for skill.
Anton Korinek (2025) · 5 citations
This paper explains what LLM-based AI agents are and gives economists a practical, code-backed playbook for using and building them to automate literature review, coding, and data analysis—while emphasizing key failure modes that require human oversight.
Francesco A. Fabozzi and Marcos López de Prado (2025) — The Journal of Portfolio Management
A practitioner-oriented guide to deploying LLMs in asset management that emphasizes governance, training, workflow integration, and the operational risk of non-reproducible (time-fragile) outputs.
Unknown
This article from Anthropic discusses best practices for building LLM agents, emphasizing simple, composable patterns over complex frameworks, and provides guidance on when to use workflows versus autonomous agents.
Unknown
The MAST-Data dataset contains execution traces of Multi-Agent Systems (MAS) annotated with the Multi-Agent Systems Failure Taxonomy (MAST), providing insights into LLM-driven agent failures.
Volodymyr Mnih et al. (2013) — arXiv:1312.5602 [cs] · 13518 citations
The paper shows that a single convolutional neural network trained with a stabilised form of Q-learning can learn to play multiple Atari 2600 games directly from pixels, achieving state-of-the-art results and even beating human experts on some games.
Zura Kakushadze et al. (2015) · 10 citations
A catalog of 101 explicit, executable quantitative trading alpha formulas with empirical analysis showing returns scale with volatility (exponent ~0.76) but are independent of turnover.
Tianqi Chen and Carlos Guestrin (2016) — Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD '16 · 51284 citations
This paper introduces XGBoost, an end-to-end gradient-boosted decision tree system that achieves state-of-the-art accuracy while scaling efficiently from a single machine to out-of-core and distributed settings via algorithmic and systems optimizations.
Diederik P. Kingma and Jimmy Ba (2017) · 164308 citations
The paper introduces Adam, a widely-used stochastic gradient optimizer that combines momentum and adaptive per-parameter learning rates with bias correction, giving robust performance on noisy, non-stationary, and sparse-gradient problems.
Ashish Vaswani et al. (2017) — arXiv:1706.03762 [cs] · 171159 citations
This paper introduces the Transformer, a sequence-to-sequence model that replaces recurrence and convolutions with multi-head self-attention, achieving state-of-the-art translation quality with much faster, highly parallel training.
Guolin Ke et al. (2017) — Curran Associates, Inc. · 13841 citations
LightGBM speeds up Gradient Boosting Decision Trees—especially on large, sparse, high-dimensional data—using gradient-aware row sampling (GOSS) and sparse-feature bundling (EFB) while keeping accuracy nearly unchanged.
Alexis Conneau and Douwe Kiela (2018) — European Language Resources Association (ELRA) · 677 citations
SentEval is a standardized, easy-to-use toolkit that evaluates universal sentence embeddings across a community-agreed suite of transfer tasks, reducing pipeline variability and making results more comparable.
Nils Reimers and Iryna Gurevych (2019) — arXiv:1908.10084 [cs] · 16856 citations
SBERT modifies BERT into a siamese/triplet architecture to produce cosine-comparable sentence embeddings that make semantic search and clustering practical (seconds instead of hours) while retaining strong accuracy on STS-style tasks.
Alejandro Barredo Arrieta et al. (2019) · 8082 citations
A comprehensive survey that defines XAI around the target audience, organizes ~400 XAI methods into taxonomies (transparent vs post-hoc; plus a deep-learning-specific taxonomy), and argues XAI must be integrated with fairness, privacy, and accountability to enable “Responsible AI.”
Yinhan Liu et al. (2019) — arXiv:1907.11692 [cs] · 28968 citations
RoBERTa shows that many post-BERT gains come from a better training recipe (more data, longer training, larger batches, no NSP, dynamic masking) rather than fundamentally new pretraining objectives, yielding state-of-the-art results on GLUE, SQuAD, and RACE.
Jacob Devlin et al. (2019) — Association for Computational Linguistics · 112230 citations
BERT introduces a Transformer encoder pre-trained with masked-token prediction (and next-sentence prediction) to learn deep bidirectional language representations that can be fine-tuned with minimal task-specific changes to achieve state-of-the-art NLP results.
Amirsina Torfi and Edward A. Fox (2020) · 79 citations
CorGAN generates privacy-preserving synthetic EHR-like records by using 1D CNN-based GANs plus a convolutional autoencoder decoder to better capture feature correlations than MLP-based baselines like medGAN.
Florian Eckerli and Joerg Osterrieder (2021) · 61 citations
This overview explains how GANs can generate statistically realistic synthetic financial data (especially time series) to mitigate data scarcity, and it includes a proof-of-concept test showing that off-the-shelf GANs can reproduce several financial “stylized facts” on S&P 500 log-returns.
Saurabh Arora and Prashant Doshi (2021) — Artificial Intelligence · 750 citations
This survey organizes inverse reinforcement learning (IRL)—learning an agent’s reward function from demonstrations—around its core technical challenges and the major method families (margin, entropy, Bayesian, classification/regression), helping you choose an IRL approach that matches your data and modeling constraints.
Gowthami Somepalli et al. (2021) · 469 citations
SAINT is a transformer-style neural network for tabular data that adds row-wise (intersample) attention and contrastive self-supervised pretraining, and it reports average benchmark performance that matches or exceeds common boosted-tree baselines.
Shuohang Wang et al. (2021) · 314 citations
The paper evaluates GPT-3 as a cheap “pseudo-labeler” to annotate training data for smaller deployable NLP models, showing large labeling cost savings (50%–96%) and further gains when mixing GPT-3 and human labels under a fixed budget.
Susan Athey et al. (2021) — Journal of Econometrics · 101 citations
The paper proposes using Wasserstein GANs to generate realistic, data-anchored Monte Carlo simulations—reducing researcher discretion in simulation design and enabling finite-sample evaluation of causal estimators on distributions that mimic a target dataset.
Zihao Zhang and Stefan Zohren (2021) · 28 citations
This paper introduces Sequence-to-Sequence and Attention-based deep learning models for multi-step Limit Order Book forecasting, demonstrating superior long-horizon accuracy and achieving ~6.5x faster training speeds using Graphcore IPUs compared to GPUs.
Campbell R. Harvey (2021) — SSRN Electronic Journal · 2 citations
A foundational overview distinguishing systematic from discretionary investing, outlining the historical evolution of algos, and warning against the dangers of overfitting in the ML era.
Xiao-Yang Liu et al. (2022) — Advances in Neural Information Processing Systems · 105 citations
FinRL-Meta is an open-source, DataOps-style pipeline + library that turns multiple real-market data sources into gym-style trading environments and reproducible DRL benchmarks (with cloud-based visualization/competitions) to reduce irreproducibility and the sim-to-real gap in financial RL.
Andrew W. Lo et al. (2023) — The Journal of Portfolio Management · 20 citations
This article surveys how NLP evolved from rule-based systems to transformer LLMs and explains what that evolution means for practical financial workflows like trading signals, risk management, and ESG/impact investing.
Azul Garza and Max Mergenthaler-Canseco (2023) · 214 citations
TimeGPT proposes a large pre-trained Transformer “foundation model” for time-series forecasting that can be used zero-shot on new datasets, aiming to match or beat strong baselines while being dramatically faster and simpler to deploy.
Qianqian Xie et al. (2023) — arXiv preprint arXiv:2306.05443 · 259 citations
PIXIU open-sources an instruction-tuned finance LLM (FinMA), a 128,640-sample finance instruction dataset (FIT), and a benchmark (FLARE) spanning NLP and prediction tasks to systematically train and evaluate LLMs for finance.
P. J. G. Lisboa et al. (2023) — Neurocomputing · 74 citations
This tutorial paper explains why interpretability/explainability has become a practical (and regulatory) necessity, clarifies post-hoc XAI vs ante-hoc interpretable modeling, and surveys methods and evaluation criteria relevant to deploying trustworthy ML in high-stakes domains.
Kezhi Kong et al. (2023) — PMLR · 72 citations
GOAT is a graph transformer that approximates global self-attention so it can scale to million-node graphs and perform well on both homophilous and heterophilous node classification.
Emre Kıcıman et al. (2023) · 426 citations
This paper evaluates whether GPT-3.5/4 can do useful causal reasoning (discovery, counterfactuals, actual causality) and finds they set new benchmark highs—while still failing unpredictably—suggesting they’re best used as domain-knowledge assistants alongside formal causal tools.
Qingsong Wen et al. (2023) · 1332 citations
A comprehensive survey of how Transformer architectures are adapted for time series tasks (forecasting, anomaly detection, classification), plus empirical diagnostics showing where time-series Transformers struggle (long inputs, deeper models) and what consistently helps (seasonal-trend decomposition).
Satyam Kumar et al. (2023) · 6 citations
A survey of 37 papers (1992–2023) showing how causal inference (e.g., DAGs/do-calculus, Bayesian networks, Granger causality, counterfactuals) is being used across banking, finance, and insurance—and arguing that banking and insurance applications are still early-stage, leaving room for impactful research and practice.
Qianqian Xie et al. (2024) — Advances in Neural Information Processing Systems · 123 citations
FinBen is a large open-source benchmark (42 datasets, 24 tasks) designed to measure what today’s LLMs can and cannot do in real financial NLP, forecasting, risk, and trading/agent settings.
Yangyang Yu et al. (2024) · 112 citations
FINCON is a GPT-4-based manager–analyst multi-agent trading system that uses CVaR-based real-time risk alerts plus “conceptual verbal reinforcement” across training episodes to improve sequential trading decisions for both single stocks and small portfolios.
Alexander Nikitin et al. (2024) · 23 citations
TSGM is an open-source framework that standardizes how to generate and evaluate synthetic time series using a wide range of generative, simulation-based, and augmentation methods, with a broad metric suite covering utility, similarity, privacy, fairness, and diversity.
Yaxuan Kong et al. (2024) — The Journal of Portfolio Management · 8 citations
This survey maps the emerging landscape of large language models (LLMs) in finance—what specialized models exist, where they can improve investment workflows, and which technical, evaluation, and ethical risks must be managed for responsible deployment.
Ruijie Tang (2024)
The paper tests whether time-series causal discovery (tsFCI, VarLiNGAM, TiMINo) can be turned into a daily long-short equity strategy, finding it can help returns but is often computationally infeasible at realistic universe sizes—except for VarLiNGAM.
Chanyeol Choi et al. (2025) · 16 citations
FINDER is a 5,703-example expert-annotated dataset of real, ambiguous finance search queries grounded in S&P 500 10‑K evidence, designed to realistically benchmark retrieval-augmented generation (RAG) for financial question answering.
Maxime Rivest (2025)
This article provides a quick introduction to DSPy, a framework for building LLM-powered applications, demonstrating how to define a task, fetch data, create a gold set, and optimize a weaker model to match a stronger model's performance.
Hugo Bowne-Anderson (2025)
This article argues against the overuse of AI agents, advocating for simpler and more structured LLM workflows, especially in enterprise settings, and highlights when agents are appropriate, such as in human-in-the-loop scenarios.
Parshin Shojaee et al. (2025) · 314 citations
The paper shows that “thinking” (chain-of-thought) reasoning models improve on medium-complexity puzzles but still hit sharp complexity thresholds where accuracy collapses—and their “thinking effort” paradoxically decreases right near failure.
Nikolaos Pippas et al. (2025) — ACM Computing Surveys · 20 citations
A structured survey of 167 papers (1996–2022) explaining how reinforcement learning is used across trading, portfolio management, execution, and market making—and why evaluation, realism, and interpretability remain the main blockers to production use.
Alejandro Lopez-Lira (2025) · 11 citations
The paper builds an open-source, realistic limit-order-book stock-market simulator where LLMs trade as heterogeneous agents, showing they can follow instructed strategies and generate market phenomena like price discovery and bubbles—useful for testing financial theories and systemic-risk scenarios.
Hoyoung Lee et al. (2025) · 9 citations
This paper proposes a way to measure and benchmark latent investment preferences (biases) in LLMs, showing that many models systematically favor certain equity styles (e.g., tech/large-cap/contrarian) and often display confirmation bias under “knowledge conflict” scenarios.
Xueying Ding et al. (2025) · 2 citations
Delphyne is a transformer time-series foundation model pre-trained on LOTSA plus finance data that argues cross-domain pre-training causes negative transfer in zero-shot, so the real value is fast few-step fine-tuning—especially for financial forecasting and risk tasks.
Matthieu Boileau et al. (2025) · 1 citations
The paper shows that basic time-series interpretation (trend, noise, extrema location) can be distilled from a large multimodal model into very small instruction-tuned LMs, enabling lightweight, natural-language explanations of temporal patterns.
Rohan Alur et al. (2025) · 1 citations
A multi-agent LLM system achieves human superforecaster performance by combining agentic search, supervisor-based reconciliation, and statistical calibration.
Alexander Rudin et al. (2025) — The Journal of Portfolio Management · 1 citations
A strategic framework for modern asset management that advocates combining economic theory (priors) with machine learning (adaptation) and human discretion to navigate market complexity.
What AI Can (and Can't Yet) Do for Alpha
Ziang Fang and Jason Moore (2025)
Man Numeric details 'AlphaGPT', a proprietary multi-agent AI workflow that automates the quantitative research cycle (ideation, coding, evaluation) to overcome human bandwidth constraints.
Alejandro Lopez-Lira et al. (2025)
LLMs can appear to “forecast” economics/markets pre-cutoff because they often memorize realized historical outcomes, making pre-cutoff forecast evaluation fundamentally non-identifiable and prone to lookahead bias.
Matt Lutey (2025) — The Journal of Portfolio Management
The paper is an instructional tutorial showing how to convert financial price time series into images (e.g., Gramian Angular Fields and Recurrence Plots) so CNNs can learn chart-like patterns for forecasting, screening, and risk monitoring.
Takaya Sekine et al. (2025) — The Journal of Portfolio Management
A practitioner tutorial showing how modern NLP (QA pipelines, news narrative construction, and ESG-claim classification) can turn unstructured text into scalable decision-support inputs for asset management.
Frank J. Fabozzi and Caleb C. Stenholm (2025) — The Journal of Portfolio Management
This paper establishes a comprehensive operational framework for asset managers by mapping military doctrines—specifically the OODA loop, mission command, and after-action reviews—to investment decision-making processes.
Guido Baltussen et al. (2025) — The Journal of Portfolio Management
This practitioner-oriented paper explains how asset managers can convert large-scale text (filings, earnings calls, news) into investable signals using NLP—from bag-of-words to LLM embeddings—while managing leakage, bias, hallucinations, translation, and cost.
The Elements of Quantitative Investing
Giuseppe A. Paleologo (2025) — John Wiley & Sons
Unknown (2025)
CrewAI OSS v1.0 is released, marking a significant milestone in agentic automation, powering 1.4 billion agentic automations and used by 60% of the Fortune 500, offering a stable, open-source framework for building complex multi-agent systems.
Hoogkamer
Termboard is a user-friendly, browser-based tool that simplifies the creation of knowledge graphs and semantic models for various applications, requiring no specialized expertise or complex installations.
Man Group
Agentic AI systems integrated with proprietary tools significantly outperform off-the-shelf LLMs in automating the implementation phase of quantitative research.
Wim De Mulder et al. — Computer Speech & Language · 259 citations
This survey reviews how recurrent neural networks (RNNs) are used for statistical language modeling, why vanilla RNN LMs work well but train slowly and struggle with long context, and which architectural/training/decoding extensions (classes, LSTM, bidirectionality, etc.) appear most promising based on reported perplexity/WER/BLEU results.
AlphaPortfolio: Direct Construction Through Deep Reinforcement Learning and Interpretable AI
Lin William Cong et al. · 60 citations
The paper proposes AlphaPortfolio, a deep offline reinforcement learning framework that directly maximizes portfolio objectives (e.g., out-of-sample Sharpe) and reports Sharpe ratios around/above 2.0 (monthly rebalancing) with >13.5% annualized factor-model alpha, plus an “economic distillation” method to interpret the learned strategy.
Deep Reinforcement Learning and Electronic Market Making
Chenyu Liu
This thesis builds a model-free deep reinforcement learning (DRL) market maker using real Bitcoin limit order book data and shows that an Advantage Actor-Critic agent can achieve positive average PnL and lower tail risk than random quoting.