Explain why point-in-time correctness and entity consistency are the core engineering constraints for fundamental and alternative data.
Implement bitemporal storage and as-of query patterns for revision-prone financial datasets.
Build a point-in-time corporate fundamentals pipeline from SEC EDGAR and XBRL filing histories.
Design time-valid entity, security, and contract mapping workflows using deterministic, probabilistic, and embedding-based resolution methods with appropriate QA gates.
Apply point-in-time alignment rules to macro, commodity, and on-chain datasets, including release timestamps, vintages, contract mapping, and finality policies.
Evaluate alternative datasets for incremental signal, data quality, legal and compliance risk, and commercial or engineering feasibility.
Extract, clean, and store SEC filing text as an auditable point-in-time corpus for downstream NLP feature engineering.
4.1
The point-in-time pipeline
The section operationalizes point-in-time correctness for fundamental data by building a bitemporal pipeline from SEC EDGAR filings. The core challenge is that amended filings, taxonomy changes, and corporate actions create multiple versions of the same reporting period — a company restating Q3 EPS months later creates lookahead bias if the restated value is used at the original date. The reader learns as-of query logic for reconstructing what was known at any historical decision time, and authoritative timestamp conventions for SEC, macro, commodity, and crypto data.
4.2
Entity resolution and mapping
A three-stage hierarchical approach to entity resolution: deterministic matching on standard identifiers (LEI, CIK, FIGI, CUSIP/ISIN), probabilistic matching using string similarity algorithms for sources lacking identifiers, and embedding-based semantic matching for complex cases like subsidiary-to-parent linking. A false-positive match poisons every downstream join, making precision more important than recall. The reader learns to build a temporally-aware master security database with time-valid identifier mappings and ongoing QA metrics.
4.3
Fundamentals across the asset-class spectrum
Point-in-time engineering is generalized beyond equities to macro/sovereign data, commodities, and crypto. Each asset class has revision-prone fundamentals with distinct timestamp authorities: macro data uses FRED's ALFRED vintage system for revision histories, commodity releases (EIA, USDA WASDE, CFTC COT) must be mapped to the correct tradable contract using consistent roll logic, and crypto on-chain fundamentals (active addresses, hash rate, TVL) have unique PIT requirements around block finality.
4.4
Understanding alternative data
A structured evaluation framework for alternative data acquisition, organized around signal content (uniqueness, decay, incrementality), data quality (versioning, coverage, latency), legal risks as a hard-fail gate (MNPI screening, privacy compliance), and commercial costs. Published return predictors lose ~58% of performance post-publication (McLean and Pontiff 2016), establishing a decay baseline. The key decision rule: privilege hard constraints over marginal score differences and only compare datasets that clear hard-fail gates.
4.5
Using text data for NLP features
The engineering pipeline for extracting model-ready text from SEC 10-K filings, focusing on MD&A (Item 7) and Risk Factors (Item 1A) as sections containing qualitative information beyond accounting line items. The four-step pipeline covers document selection, section extraction with quality checks, cleaning that preserves paragraph structure, and PIT-correct storage using accession numbers with both filing dates and SEC acceptance timestamps. This corpus serves as input for Chapter 10's NLP feature engineering.
4.6
Summary
01 Academic Characteristics
02 Sec Filing Explorer
03 Sec Form4 Insider Transactions
04 Sec Xbrl Fundamentals
05 Entity Resolution
06 Fred Macro Eda
07 Macro Data Alignment
08 Futures Positioning
09 Onchain Fundamentals
10 Institutional Holdings 13F
11 Defi Tvl Evaluation
12 Kalshi Prediction Markets
13 Polymarket Prediction Markets
14 Text Data Extraction
5 primer topics providing foundational concepts for this chapter.
Harrison Hong et al. (2000) — The Journal of Finance
Momentum profits are significantly higher in stocks with low analyst coverage, particularly for past losers, supporting the theory that information (especially bad news) diffuses gradually.
Kent Daniel and Sheridan Titman (2006) — The Journal of Finance · 206 citations
The paper decomposes stock returns into 'tangible' (accounting-based) and 'intangible' components, finding that the value premium and long-term reversals are driven entirely by overreaction to intangible information, while the market reacts efficiently to tangible fundamental performance.
This paper reviews research on real-time data analysis, focusing on the impact of data revisions on forecasting, monetary policy, macroeconomic research, and current analysis of business and financial conditions, highlighting progress and areas needing further exploration.
This seminal paper demonstrates that increases in Google search volume for financially negative terms like 'debt' precede stock market falls, enabling a strategy that returned 326% over 7 years.
Seven Sins of Quantitative Investing
Yin Luo et al. (2014)
A comprehensive empirical audit of seven common backtesting biases, demonstrating how errors in data handling (survivorship, look-ahead) and modeling (outliers, signal decay) can invert strategy performance from profitable to disastrous.
Paul C. Tetlock (2014) — Annual Review of Financial Economics
A comprehensive review of how media content, linguistic tone, and investor attention influence asset prices, establishing the foundation for Natural Language Processing (NLP) in finance.
Point-In-Time vs. Lagged Fundamentals
Ernest Breitschwerdt (2015)
This paper argues that using lagged non-point-in-time (Non-PIT) financial data in backtesting can lead to significantly different and potentially misleading results compared to using point-in-time (PIT) data, due to varying filing regulations, restatements, and the overwriting of historical data.
Rethinking Alternative Data in Institutional Investment
Ashby Monk et al. (2019)
This paper argues that institutional investors should focus on defensive and defensible strategies for using alternative data, emphasizing risk management and operational alpha, rather than solely pursuing alpha-generating opportunities.
Gene Ekster and Petter N. Kolm (2020) · 5 citations
A comprehensive guide to the alternative data ecosystem, detailing a specific preprocessing pipeline (entity tagging, stabilization, debiasing) that reduces revenue prediction error from 88% to 2.6%.
This paper introduces a deep learning model that uses a no-arbitrage condition to estimate a flexible, time-varying stochastic discount factor (SDF) for individual stock returns, significantly outperforming traditional and other machine learning benchmarks in explaining cross-sectional returns and identifying key risk drivers.
Alfred Lehar and Christine A. Parlour (2021) · 83 citations
This paper analyzes Uniswap, a prominent decentralized exchange using automated market making (AMM), characterizing equilibrium in liquidity pools, providing empirical evidence of their stability, and comparing it to centralized exchanges like Binance.
Dirk G. Baur and Lee A. Smales (2022) — Journal of Futures Markets
Leveraged money traders (hedge funds) in Bitcoin futures act as 'smart money' by successfully timing the market through adjustments in their short positions, offering a profitable signal for other investors.
A comprehensive analysis of crypto asset characteristics demonstrating that while fundamental valuation models are flawed, active strategies like volatility targeting and trend following significantly improve risk-adjusted returns.
Florian Berg et al. (2022) — Review of Finance · 1489 citations
ESG ratings from major providers exhibit low correlation (average 0.54), driven primarily by disagreements on how to measure specific attributes (56% of divergence) rather than what attributes to include or how to weight them.
Natthawut Kertkeidkachorn et al. (2023) — 2023 IEEE 17th International Conference on Semantic Computing (ICSC) · 14 citations
This paper introduces FinKG, a high-quality financial knowledge graph with a manually crafted ontology, demonstrating its utility for complex knowledge retrieval and improving stock price prediction by providing rich, structured financial features.
This paper demonstrates that sell-side financial analysts are increasingly adopting alternative data, which significantly improves their earnings forecast accuracy and increases trading commissions for their brokerages, suggesting investors value these insights and that analysts can maintain relevance in a data-driven world.
T. Clifton Green and Shaojun Zhang (2024) — The Journal of Portfolio Management
A comprehensive survey of the alternative data landscape, categorizing sources into firm-released, government, attention, and third-party data, while analyzing the regulatory drivers and implementation challenges like alpha decay.
Jacques Joubert et al. (2024) — The Journal of Portfolio Management
A practitioner-focused guide to designing more reliable backtests by choosing the right backtest type, avoiding common simulation biases, and correcting Sharpe-ratio inference for multiple strategy trials.
This paper provides the first empirical evidence of cross-market price discovery in modern prediction markets, finding significant arbitrage opportunities and demonstrating that Polymarket leads Kalshi in information aggregation, driven by liquidity and informed 'whale' trades.
Joseph D. Piotroski (2000) — Journal of Accounting Research · 1151 citations
Introduces the 'F-Score', a 9-point composite metric based on financial statements that separates winners from losers within high book-to-market portfolios, increasing annual returns by 7.5% over generic value strategies.
An update to Varian's seminal work on search data, introducing Bayesian Structural Time Series (bsts) as a superior method for selecting relevant Google Trends predictors to forecast economic indicators.
This corrigendum fixes a look-ahead/data-leakage issue in an earnings-forecasting ML setup and shows that ML still beats analysts on forecast accuracy, while the return-predictability linked to analyst optimism remains but is quantitatively weaker.
This paper examines how Apple's App Tracking Transparency (ATT) policy, which restricts cross-app data sharing, weakened the predictive power of mobile-generated data for firm performance and investment decisions, leading to information frictions and potential mispricing in financial markets.
Iro Tasitsiomi and Yijie Wang (2025) — The Journal of Portfolio Management
A practitioner's framework for 'quantamental' investing that categorizes how alternative data and AI models create new factors to augment, rather than replace, fundamental analysis.
Yosef Bonaparte and Frank J. Fabozzi (2025) — The Journal of Portfolio Management
The authors construct a 'Fear of Missing Out' (FoMO) index using social media, momentum, and margin debt data, finding it predicts asset bubbles, Bitcoin trading volume, and subsequent market corrections.
Ray Ball and Philip Brown (1968) — Journal of Accounting Research
This paper empirically demonstrates that accounting income numbers contain useful information for investors, as reflected in stock price movements, but most of this information is anticipated by the market well before the official annual report release.
Arturo Estrella and Gikas A. Hardouvelis (1991) — The Journal of Finance
The slope of the yield curve, specifically the spread between 10-year Treasury bonds and 3-month T-bills, is a robust and superior predictor of future real economic activity (GNP, consumption, investment) up to four years ahead, outperforming traditional indicators and survey forecasts.
The paper uses topic modeling on 10-K “Item 1A Risk Factors” text to infer firm-specific risk exposures, builds tradable risk-mimicking portfolios and a parsimonious factor model from them, and shows these “firm-identified risk factors” add pricing information beyond standard factor models.