ML4T Engineer
ML4T Engineer Documentation
Features, labels, alternative bars, and leakage-safe dataset preparation
Skip to content

Book Guide

This guide maps ml4t-engineer tasks to public notebooks from Machine Learning for Trading, Third Edition. Every link uses companion commit d2edec54b1c7a6a9d7a97d8129eb05db4491e1eb. All paths and the paired Python sources were checked at that commit.

The relationship column distinguishes three cases:

  • Calls Engineer: the notebook imports and runs the named ml4t.engineer API.
  • Teaches manually: the notebook implements the method for instruction and does not use Engineer for that task.
  • Related workflow: the notebook shows where the task fits, but its broader workflow is not an Engineer API example.

Feature computation and discovery

Book notebook Relationship Engineer API Task guide
The ml4t Library Ecosystem Calls Engineer to inspect registry metadata and run compute_features() with names and parameter dictionaries compute_features, get_registry Features, Feature Discovery
Price and Volume Feature Families Calls Engineer for registry features, volatility, regime, risk, and fractional-differencing functions; also derives selected features manually compute_features and feature modules Features, ML Readiness
Microstructure Features Calls Engineer for the tick rule and liquidity estimators, then builds a wider teaching workflow ml4t.engineer.features.microstructure Features
Structural and Cross-Instrument Features Calls Engineer for market beta; teaches carry and options features manually beta_to_market Features
Slow Features and Context Calls Engineer for calendar encoding; teaches point-in-time joins and slow features manually cyclical_encode Features
Panel Features Calls Engineer for cross-asset features and compares them with manual statistical work ml4t.engineer.features.cross_asset Features
ETFs: Feature Engineering Calls individual Engineer feature functions inside a full case-study pipeline momentum, trend, volatility, volume, and regime feature modules Features

The chapter notebooks use book datasets and plotting dependencies. The Quickstart provides an offline synthetic path for the same released computation API.

Labeling

Book notebook Relationship Engineer API Task guide
Label Engineering Methods Calls Engineer for fixed-horizon, percentile, triple-barrier, ATR-barrier, trend-scanning, meta-labeling, and sample-weighting workflows ml4t.engineer.labeling, LabelingConfig Labeling
ETFs: Label Engineering Related workflow that constructs and audits case-study labels without calling Engineer no direct Engineer call Labeling

The second notebook is useful for the artifact and timing workflow. It is not evidence that the case study uses Engineer's labeling functions.

Alternative bars

Book notebook Relationship Engineer API Task guide
ITCH Bar Sampling Calls Engineer for tick, volume, dollar, imbalance, and run bars on ITCH trades bar sampler classes Alternative Bars
Information-Bar Formulas and Parameters Calls Engineer and compares manual formulas with adaptive, fixed, and window samplers imbalance-bar sampler classes Alternative Bars
Databento Bar Calibration Calls Engineer in a multi-day calibration workflow that requires Databento data bar sampler classes Alternative Bars

The first two notebooks require book data. The Databento notebook also requires the vendor dataset. The task guide and examples/bars_example.py provide an offline, synthetic verification path.

Preprocessing and fractional differencing

Book notebook Relationship Engineer API Task guide
Preprocessing Pipeline Calls Engineer's StandardScaler for train-only fitting; teaches the broader cleaning pipeline manually StandardScaler Preprocessing
Fractional Differencing Calls Engineer's fractional-differencing helpers while teaching the statistical method ffdiff, find_optimal_d, fdiff_diagnostics Fractional Differencing
ETFs: Model-Based Features Calls ffdiff inside a walk-forward case-study workflow; its HMM and GARCH work is outside Engineer's fractional-differencing API ffdiff Fractional Differencing

No checked book notebook at this revision calls create_dataset_builder. Use the Dataset Builder guide and the repository's examples/complete_workflow_example.py for that workflow.

Alphalens migration scope

Engineer overlaps with Alphalens only before factor analysis: it can compute factor values with compute_features() and fit preprocessing state on training data. Engineer does not replace Alphalens tearsheets, information-coefficient analysis, quantile-return analysis, turnover analysis, or event studies. Use ML4T Diagnostic for those evaluation tasks. The book's feature notebooks above show factor construction; topical similarity does not make them Alphalens replacements.

Supported and experimental boundaries

  • The principal documented workflows are feature computation and discovery, labeling, alternative bars, preprocessing, dataset building, and fractional differencing.
  • Cross-asset functions are supported advanced APIs. Validate asset ordering and point-in-time alignment before use.
  • Adaptive imbalance bars require calibration. Fixed-threshold samplers provide the bounded production path described in the Alternative Bars guide.
  • fdiff_diagnostics() and find_optimal_d() require the stats extra. ffdiff() is available from the core installation.
  • The optional DuckDB store is experimental and has no verified book adoption path.
  • transfer_entropy() is not implemented for production use. It is not part of the principal documented workflow.

Run a workflow first