Apply a three-timestamp framework and disclosure-time cutoff rules to prevent temporal leakage in graph queries,
Make sound engineering choices about graph databases, ontology scope, query safety, and schema evolution for
23.1
When relational structure justifies a graph
This section establishes a decision framework for when graph infrastructure is justified, identifying three conditions that reliably add value: multi-hop dependency queries (supply chain contagion analysis), structural crowding and co-ownership patterns (institutional holdings that predict excess comovement), and temporal relationship evolution (edges that change before price data reflects the shift). Equally important, it identifies three conditions where graphs do not help: single-entity attribute lookups, narrative synthesis over broad corpora, and sparse graphs with few relationships where topology metrics become unreliable. The practical test is whether the question naturally decomposes into path patterns across entities.
23.2
Constructing financial knowledge graphs
The section presents a governance-first approach to KG construction where identity contracts, schema contracts, and provenance contracts are defined before any extraction begins, then paired with LLM-based extraction for throughput. A five-stage workflow covers targeted document slicing, schema-constrained generation, canonicalization, rule validation, and human review queuing, with emphasis on idempotent writes and quality metrics beyond precision (schema-valid rate, provenance coverage, duplicate-node rate, temporal consistency). The supply chain schema example demonstrates extracting supplier, customer, and competitive relationships from S&P 100 10-K filings, revealing shared suppliers like TSMC and Foxconn as concentration nodes.
23.3
Deterministic relational reasoning with graph RAG
Graph RAG delegates relational joins to the database engine and language generation to the LLM, providing deterministic multi-hop retrieval that vector search cannot match for path-dependent questions. The five-stage architecture covers query routing, text-to-Cypher generation with safety validation, deterministic execution, and grounded synthesis with two-layer citations linking graph rows to underlying disclosure text. FinReflectKG-MultiHop benchmark evidence shows KG-guided retrieval improving correctness by approximately 24% while reducing token consumption by roughly 85% compared to page-window retrieval, demonstrating that structured retrieval is both more accurate and more economical for relational questions.
23.4
From graphs to machine learning features
This section transforms graph structure into tabular features for gradient boosting and factor models, covering network topology features (PageRank, betweenness, clustering coefficient), supply chain risk indicators (supplier count, single-source dependency, overlap ratios), and institutional crowding features (crowding score, smart money concentration, ownership HHI, co-ownership Jaccard). The distinctive value emerges from cross-graph integration: combining supply chain concentration with institutional crowding creates compound risk features that encode structural dependencies traditional factor models treat as independent. Temporal dynamics such as relationship churn, centrality momentum, and event propagation lags add lead-lag features that capture evolving structure invisible to static snapshots.
23.5
Correlations and portfolios in financial networks
The section connects the established financial networks literature (correlation-based MSTs, hierarchical clustering) to knowledge graph features, showing how network-aware portfolio allocation weights peripheral stocks more heavily for diversification benefits. It provides a maturity assessment of graph neural networks across financial domains: fraud detection is production-ready with clear labels and high signal-to-noise ratios, systemic risk monitoring is emerging in regulatory pilots, and alpha generation remains experimental with limited reproducible evidence after costs. The pragmatic recommendation is to start with hand-crafted graph features as the auditable, stable baseline and add GNNs only if they improve out-of-sample metrics after transaction costs and temporal leakage controls.
23.6
Temporal integrity and leakage-safe evaluation
This section introduces the three-timestamp model (event time, disclosure time, extraction time) as central to financial knowledge graphs, arguing that feature generation and historical queries must use disclosure time as the visibility gate since economically true but undisclosed information cannot enter models. Using 8-K filings and 13F holdings as case studies, it demonstrates how disclosure lags, position masking, and confidential treatment gaps create temporal inconsistencies that must be handled explicitly. The leakage-safe evaluation protocol requires splitting by disclosure time, embargoing observations near split boundaries, enforcing cutoff filtering in every query path, and logging snapshot hashes for replay and audit.
23.7
Building a KG-ready pipeline
The section addresses the infrastructure choices that make knowledge graph workflows reliable: database engine selection (Neo4j as the teaching default, with PostgreSQL recursive CTEs as a simpler alternative for narrow graphs), ontology strategy (starting compact with 8-15 relationship types rather than attempting full FIBO adoption), and query safety controls including read-only credentials, allowlisted labels, parameterized queries, and execution limits. Schema versioning discipline is emphasized, with additive changes as the safe default and version metadata on edges enabling coexistence of old and new extraction formats during transitions.
23.8
Summary
01 Sp100 Sec Download
02 Supply Chain Kg Construction
02 Supply Chain Kg Construction Qwen25 Rerun
02 Supply Chain Kg Construction Qwen3
03 Graph Rag Qa
04 Rag Comparison Benchmark
05 Institutional Holdings Kg
06 Gnn Feature Engineering
07 Dynamic Kg Temporal
08 8K Event Extraction
08 8K Event Extraction Qwen3
09 Knowledge Graph Features
10 Network Portfolio Construction
2 primer topics providing foundational concepts for this chapter.
Christopher J. C. H. Watkins and Peter Dayan (1992) — Machine Learning · 12139 citations
The paper proves that tabular Q-learning converges almost surely to the optimal action-value function in finite Markov decision processes under standard step-size and exploration conditions.
R. N. Mantegna (1999) — Computer Physics Communications · 47 citations
The paper shows how to extract economically meaningful “sector-like” structure from stock price time series by converting correlations into distances and building a minimum spanning tree (MST) / ultrametric hierarchy.
The paper proves that policy-gradient (actor-critic) reinforcement learning can be made theoretically well-behaved with function approximation by using a “compatible” value/advantage approximator, yielding convergence to a locally optimal policy.
Andrew G. Haldane and Robert M. May (2011) — Nature · 1314 citations
This paper uses ecological models to analyze systemic risk in financial networks, arguing that increased complexity and homogeneity in banking can lead to instability and proposing policy interventions like higher capital requirements and shaping network topology.
The paper shows that a single convolutional neural network trained with a stabilised form of Q-learning can learn to play multiple Atari 2600 games directly from pixels, achieving state-of-the-art results and even beating human experts on some games.
Volodymyr Mnih et al. (2015) — Nature · 30995 citations
This paper introduces the Deep Q-Network (DQN), showing that a single deep reinforcement learning algorithm can learn directly from pixels to reach human-comparable performance across dozens of Atari 2600 games.
Thomas N. Kipf and Max Welling (2017) · 34181 citations
This paper introduces Graph Convolutional Networks (GCNs), a scalable approach for semi-supervised learning on graph-structured data, achieving state-of-the-art results in node classification tasks on citation networks and knowledge graphs.
This paper introduces Graph Attention Networks (GATs), a novel neural network architecture for graph-structured data that uses masked self-attention to weigh the importance of different nodes in a neighborhood, achieving state-of-the-art results on several graph benchmarks.
William L. Hamilton et al. (2018) · 18972 citations
This paper introduces GraphSAGE, a novel inductive node embedding approach that leverages node features to generate embeddings for unseen nodes in evolving graphs, outperforming transductive methods and offering generalization across different graphs.
Petter N. Kolm and Gordon Ritter (2019) · 69 citations
This paper explains how to cast classic intertemporal finance problems (trading, hedging, execution, portfolio choice) as reinforcement learning (RL) tasks by designing reward functions that approximate expected utility, and illustrates the approach with simulated mean-reversion trading and cost-aware option hedging.
This paper introduces a novel dynamic financial knowledge graph for A-share listed companies, utilizing transfer learning and reinforcement learning to extract and classify financial entities and relationships from multi-source data, visualized through an interactive system.
This paper introduces the Elliptic Data Set, a large labeled graph network of Bitcoin transactions, and explores machine learning methods, including Graph Convolutional Networks (GCNs), for anti-money laundering (AML) to balance financial safety with inclusion.
The paper trains deep reinforcement learning agents to output futures trading positions directly (rather than forecasting returns), and shows they beat classic time-series momentum baselines on 50 liquid futures from 2011–2019—even with large transaction costs.
Sarah Elhammadi et al. (2020) — International Committee on Computational Linguistics · 40 citations
This paper introduces a high-precision pipeline for automatically constructing a financial knowledge graph (KG) by extracting information from financial news articles, achieving 78% precision on the top 100 extractions.
Dawei Cheng et al. (2020) — Association for Computing Machinery · 89 citations
This paper introduces a knowledge graph-based event embedding framework (KGEEF) to improve quantitative investment strategies by incorporating lead-lag relationships between entities extracted from financial news.
This paper introduces Retrieval-Augmented Generation (RAG), a novel approach that combines pre-trained sequence-to-sequence models with a non-parametric memory to improve performance on knowledge-intensive NLP tasks.
Thibaut Théate and Damien Ernst (2021) — Expert Systems with Applications · 205 citations
The paper proposes a Deep Q-Learning-based trading agent (TDQN) that learns daily long/short positioning from OHLCV history and evaluates it with a more rigorous multi-stock testbench, showing promising (but variance-prone) Sharpe-ratio performance versus classic rules.
Gautier Marti et al. (2021) — Springer International Publishing · 137 citations
A comprehensive review of correlation-based clustering and network methods (MST, PMFG, hierarchical clustering, RMT, etc.) for understanding market structure, building portfolios, and monitoring systemic risk—along with why these methods are often unstable and hard to operationalize.
This paper proposes an automated querying engine using a Financial Knowledge Graph (FKG) and Ontology to extract information from annual financial reports, enabling stakeholders to make informed decisions and generate custom financial stories.
A critical survey of deep reinforcement learning (DRL) for trading—especially crypto—mapping common design patterns (state/action/reward) and highlighting why inconsistent datasets, environments, and cost modeling make results hard to compare or trust.
This survey organizes inverse reinforcement learning (IRL)—learning an agent’s reward function from demonstrations—around its core technical challenges and the major method families (margin, entropy, Bayesian, classification/regression), helping you choose an IRL approach that matches your data and modeling constraints.
The paper builds a realistic-ish stock-trading reinforcement-learning environment (with technical indicators, FinBERT headline sentiment, and transaction costs) and shows a TD3 agent can reach a 2.68 Sharpe ratio on an out-of-sample 10-stock test period.
Charles K. Assaad et al. (2022) — Journal of Artificial Intelligence Research · 184 citations
A comprehensive survey of causal discovery methods for multivariate time series, plus an empirical comparison showing that performance is highly assumption- and dataset-dependent—no single method dominates.
Andrew Lampinen et al. (2023) — Advances in Neural Information Processing Systems · 25 citations
The paper shows that agents—and even pretrained language models—can learn generalizable “do experiments, infer causal structure, then exploit” strategies from purely passive data, as long as they can intervene at test time, especially when training includes natural-language explanations.
The paper designs a DDPG-based RL hedger for short-dated equity index options that explicitly estimates decision uncertainty (aleatoric + epistemic) to reduce hedging P&L variance and avoid overconfident, high-turnover re-hedging.
Natthawut Kertkeidkachorn et al. (2023) · 14 citations
This paper introduces FinKG, a high-quality financial knowledge graph with a manually crafted ontology, demonstrating its utility for complex knowledge retrieval and improving stock price prediction by providing rich, structured financial features.
Gueorgui S. Konstantinov et al. (2023) — The Journal of Portfolio Management · 1 citations
This paper is a non-technical guide to using network (graph) representations of asset relationships to improve portfolio diversification decisions and systemic-risk awareness beyond what correlations and regressions show.
Shuo Sun et al. (2023) — ACM Transactions on Intelligent Systems and Technology · 81 citations
A 2023 survey that organizes and critiques RL methods across algorithmic trading, portfolio management, order execution, and market making, and argues the field needs more realistic simulators and more unified evaluation to make RL trading research actionable.
Frank J. Fabozzi et al. (2024) — The Journal of Portfolio Management
The paper argues that causal modeling for investing should combine (1) empirical data, (2) theory-implied causal structure, and (3) real-time qualitative “priors” into one structural causal modeling workflow to better navigate non-mechanistic financial systems.
Xiaohui Victor Li and Francesco Sanna Passino (2024) — Association for Computing Machinery · 21 citations
This paper introduces a novel framework for financial trend detection using dynamic knowledge graphs (DKGs) generated by a fine-tuned Large Language Model (ICKG) from financial news, which are then analyzed by a new graph neural network (KGTransformer) to outperform existing thematic investing strategies.
This paper provides the first comprehensive survey of Graph Retrieval-Augmented Generation (GraphRAG), a novel approach that enhances Large Language Models by leveraging structured relational knowledge from graph databases to mitigate issues like hallucination and lack of domain-specific context.
Gueorgui S. Konstantinov and Frank J. Fabozzi (2025) — The Journal of Portfolio Management · 1 citations
The paper proposes CAFNITE, a causal-network framework that measures how shocks to one factor/asset propagate through a global multi-asset factor network (2001–2024), showing diversification depends on time-varying causal linkages rather than static correlations and offering early-warning diagnostics before volatility spikes.
This paper introduces FinReflectKG - MultiHop, a new benchmark for multi-hop financial question answering, demonstrating that using knowledge graph (KG)-linked evidence significantly boosts LLM accuracy and efficiency compared to traditional text-window retrieval.
Abhinav Arun et al. (2025) — Association for Computing Machinery · 9 citations
This paper introduces FinReflectKG, an open-source, large-scale financial knowledge graph dataset built from S&P 100 SEC 10-K filings, alongside an agentic, reflection-driven framework for its construction and a holistic evaluation methodology, demonstrating superior extraction quality.
FINDER is a 5,703-example expert-annotated dataset of real, ambiguous finance search queries grounded in S&P 500 10‑K evidence, designed to realistically benchmark retrieval-augmented generation (RAG) for financial question answering.
This paper introduces GraphRAG, a novel graph-based retrieval augmented generation approach that enables global sensemaking over large text corpora by constructing a knowledge graph, partitioning it into a hierarchy of communities, and generating community-level summaries to answer queries.
A structured survey of 167 papers (1996–2022) explaining how reinforcement learning is used across trading, portfolio management, execution, and market making—and why evaluation, realism, and interpretability remain the main blockers to production use.
Yadh Hafsi and Edoardo Vittori (2025) · 5 citations
The paper trains a high-frequency (1-second) reinforcement learning agent in the ABIDES multi-agent limit-order-book simulator to learn an optimal execution policy that beats standard schedules like TWAP while controlling market impact.
CrewAI OSS v1.0 is released, marking a significant milestone in agentic automation, powering 1.4 billion agentic automations and used by 60% of the Fortune 500, offering a stable, open-source framework for building complex multi-agent systems.
Machine Learning for Market Microstructure and High Frequency Trading
Michael Kearns and Yuriy Nevmyvaka · 64 citations
This paper demonstrates how to apply Machine Learning (RL and censored estimation) to three core HFT problems: optimized trade execution, directional price prediction, and dark pool order routing.
Deep Reinforcement Learning and Electronic Market Making
Chenyu Liu
This thesis builds a model-free deep reinforcement learning (DRL) market maker using real Bitcoin limit order book data and shows that an Advantage Actor-Critic agent can achieve positive average PnL and lower tail risk than random quoting.
Albert S. Kyle (1985) — Econometrica · 9334 citations
The foundational paper establishing 'Kyle's Lambda' as a measure of market impact, proving that an informed monopolist trades gradually to disguise information, resulting in constant market depth and prices that follow Brownian motion.
Backpropagation through time: what it does and how to do it
P.J. Werbos (1990) — Proceedings of the IEEE · 5220 citations
The paper explains how to compute exact gradients efficiently for dynamic/recurrent systems using backpropagation through time (BPTT), with equations and pseudocode that generalize basic backprop to time-lagged networks and control/identification problems.
Yuriy Nevmyvaka et al. (2006) — Association for Computing Machinery · 291 citations
This paper shows—using 1.5 years of millisecond NASDAQ limit-order-book data—that reinforcement learning can learn execution policies that reduce implementation shortfall by up to ~50% versus optimized “submit-and-leave” baselines.
Marco Avellaneda and Sasha Stoikov (2008) — Quantitative Finance · 522 citations
A seminal mathematical framework for market making that derives optimal bid-ask quotes as a function of inventory levels and market volatility to maximize risk-adjusted returns.
M. Dai et al. (2010) — SIAM Journal on Financial Mathematics
This paper mathematically proves that a trend-following strategy based on two probability thresholds (buy/sell) is optimal in a regime-switching market and outperforms buy-and-hold by ~2x in historical tests.
Bence Toth et al. (2011) — Physical Review X · 243 citations
The paper provides a theoretical and empirical basis for the 'square-root law' of price impact, arguing that markets self-organize into a critical state where liquidity vanishes linearly near the current price.
Anna A. Obizhaeva and Jiang Wang (2013) — Journal of Financial Markets
This paper introduces a dynamic Limit Order Book (LOB) model where liquidity replenishes over time ('resilience'), proving that optimal execution requires a mix of large discrete trades at the start and end, bridged by continuous trading.
The paper tests whether inverse reinforcement learning can recover an expert trader’s (possibly non-linear) objective from limit order book demonstrations, finding that Bayesian neural network and Gaussian-process IRL succeed on non-linear rewards while linear MaxEnt IRL does not.
Chun-Chieh Wang and Yun-Cheng Tsai (2019) — arXiv:1908.08036 [cs] · 6 citations
The paper tests whether deep reinforcement learning (especially PPO) can directly learn FX trading decisions by optimizing a Martingale-like “Sure-Fire” policy using price-chart state representations (GAF images), instead of predicting prices.
Zhicheng Wang et al. (2021) — Proceedings of the AAAI Conference on Artificial Intelligence · 146 citations
DeepTrader is a deep reinforcement learning portfolio manager that explicitly embeds market conditions to dynamically shift between long and short exposure, improving risk-return balance (especially drawdowns) versus prior DRL baselines.
Ryan Donnelly (2022) — Applied Mathematical Finance · 15 citations
A comprehensive review of optimal execution models, ranging from the foundational Almgren-Chriss framework to modern extensions involving stochastic volatility, transient price impact, and alternative objectives like VWAP targeting.
This paper proposes a coarse-grained theoretical model proving that the 'square-root law' of market impact arises from a specific supply-demand equilibrium where the ratio of average impact to peak impact stabilizes at 2/3.
Andy Novocin and Bruce Weber (2022) — The Journal of Portfolio Management · 2 citations
The paper argues that blockchain, Generative AI, and DAOs will disrupt financial markets by decoupling trading functions from centralized exchanges, similar to how Wikipedia disrupted Britannica.
Xiao-Yang Liu et al. (2022) — Advances in Neural Information Processing Systems · 105 citations
FinRL-Meta is an open-source, DataOps-style pipeline + library that turns multiple real-market data sources into gym-style trading environments and reproducible DRL benchmarks (with cloud-based visualization/competitions) to reduce irreproducibility and the sim-to-real gap in financial RL.
Yongjae Lee et al. (2023) — The Journal of Portfolio Management · 18 citations
A comprehensive survey categorizing machine learning models by their specific applications in asset management, ranging from tree-based models for tabular data to reinforcement learning for hedging.
Hanchen Wang et al. (2023) — Nature · 1475 citations
A comprehensive review of how AI—specifically geometric deep learning, self-supervised learning, and generative models—is transforming the scientific process from hypothesis generation to autonomous experimentation.
Yongjae Lee et al. (2024) — The Journal of Portfolio Management · 14 citations
This educational overview explains how machine learning can improve both the inputs to portfolio optimization (returns, risk, similarity) and the optimization step itself, and how newer “decision-focused” and end-to-end methods move beyond the classic predict-then-optimize pipeline.
Petter N. Kolm and Gordon Ritter (2025) — The Journal of Portfolio Management
An equation-free, practitioner-oriented overview of how reinforcement learning (RL) can be used to learn adaptive trading, hedging, and portfolio policies under frictions where today’s actions change tomorrow’s opportunities.
Limit Order Flow, Market Impact and Optimal Order Sizes: Evidence from NASDAQ TotalView-ITCH Data
Nikolaus Hautsch · 44 citations
This paper quantifies the permanent price impact of limit orders on NASDAQ, demonstrating that limit orders (not just market orders) move prices, and provides a method to calculate optimal order sizes to minimize this impact.
This is a comprehensive textbook on the foundations of machine learning, covering supervised learning, deep learning, causality, reinforcement learning, and more, emphasizing the interplay between patterns, predictions, and actions.
The MAST-Data dataset contains execution traces of Multi-Agent Systems (MAS) annotated with the Multi-Agent Systems Failure Taxonomy (MAST), providing insights into LLM-driven agent failures.