← All courses

Self-paced · ml4trading.io

ML for Trading: Foundations

For anyone building their first end-to-end ML trading strategy, and for Research to Production students closing gaps.

Build one strategy, once, end to end: from the mechanism you claim makes money through to the point where it would be handed to a live broker, on a 100-ETF universe spanning equities, bonds, commodities, currencies and real estate, priced daily from free data. You watch the finished pipeline run once, so you have the map, and then you build it yourself: what you have at each stage is narrow and correct rather than broad and broken. The recommended route into Research to Production, and available on its own.

How it works  Each unit is a recorded segment with its own reading, notebook, exercise and check for understanding, plus a standing biweekly research hour for as long as you are enrolled.

Research hours

A standing biweekly research hour, free and live while you are enrolled. Bring what you are stuck on. Upcoming dates:

Fri, Oct 9 Fri, Oct 23 Fri, Nov 6 Fri, Nov 20

Syllabus

Eight parts, forty-two units. Each part settles one decision about the pipeline you are building, and each one widens what that pipeline handles: it is correct and narrow at every stage, never full-scope and wrong.

  1. Part 1 5 units

    Orientation and Strategy Definition

    Decision

    What you are committing to build, which economic hypothesis it tests, and how its results will be judged - all fixed before any data is touched.

    Builds

    A stated strategy hypothesis, the objective metric and transaction cost assumptions that every later comparison applies, and a feasibility assessment against history, breadth, and decision-time information.

    • The ML4T Workflow
    • The Finished Pipeline, End to End
    • Trading Strategy Families and Sources of Edge
    • Transaction Cost Assumptions and Performance Metrics
    • Feasibility and Breadth
  2. Part 2 7 units

    Data, Universe, and a Non-ML Baseline

    Decision

    Which data and instruments enter the first executable strategy, when its decisions are made, and which non-ML rule later models must beat.

    Builds

    A validated, point-in-time-clean 100-ETF panel, the eligible universe and decision clock, and a reproducible non-ML baseline run with its environment, state, seed, and run log.

    • Data Sources, Coverage, and the Daily Close
    • Point-in-Time Data and the Availability Lag
    • Adjustment Policy and Quality Gates
    • The Eligible Universe and the Decision Clock
    • The Non-ML Baseline
    • Running the Pipeline End to End
    • The Run You Can Reproduce
  3. Part 3 6 units

    Validation and Labels

    Decision

    How you will find out whether anything works, and what the model is allowed to learn from - both fixed before there is anything to evaluate.

    Builds

    The walk-forward scheme with its purge and embargo, a sealed holdout with a written rule for who may open it, an execution-consistent label set, and a preprocessing pipeline with every fitted transform inside the fold.

    • Walk-Forward Validation, Purging, and Embargo
    • The Sealed Holdout
    • Label Definition, Horizon, and Overlap
    • From Decision Time to Executable Price
    • Regression, Classification, or Quantile
    • Preprocessing Inside the Fold
  4. Part 4 6 units

    Feature Engineering

    Decision

    Which features earn a place, given that each one is a claim about a driver and each one spends a degree of freedom.

    Builds

    The engineered ETF feature panel with range-based volatility, cross-sectional and relative-value features, a justified stationarity transformation, and the degrees of freedom the feature search consumed.

    • Feature Design and Economic Rationale
    • Lookback Windows and Normalization
    • Range and Volatility Estimators
    • Cross-Sectional and Relative-Value Features
    • Stationarity Diagnostics and Feature Transformations
    • Feature Count and Degrees of Freedom
  5. Part 5 6 units

    Models: A Linear Baseline and Gradient Boosting

    Decision

    Whether added model capacity has earned its place over a simple baseline.

    Builds

    A regularized linear model and a tuned gradient-boosted model on identical folds, with the search budget spent on each recorded.

    • Regularized Linear Models
    • Running the Pipeline With a Model
    • Gradient Boosting
    • Running the Pipeline With Gradient Boosting
    • Loss Functions, Model Selection, and the Search Budget
    • Deep Learning, Latent Factor and Causal Models
  6. Part 6 4 units

    Backtesting

    Decision

    Whether the signal survives becoming a strategy under explicitly stated mechanics.

    Builds

    A backtest of the ETF strategy with its mechanics stated in full, compared with the baseline fixed before any of it was run.

    • From Score to Signal
    • Backtest Engine, Fills, and Rebalance Frequency
    • Comparing the Result With the Baseline
    • Running the Pipeline With a Backtest Engine
  7. Part 7 5 units

    Allocation, Costs, and Risk

    Decision

    Whether the strategy still works once positions are sized, costs are charged, and risk is controlled.

    Builds

    A version of the strategy with costs charged, positions allocated and risk controlled with a stated break-even turnover.

    • Mapping Scores to Weights
    • Portfolio Constraints and the Benchmark
    • The Cost Model and Its Parameters
    • Position Controls and Exits
    • Running the Pipeline With Allocation and Costs
  8. Part 8 3 units

    Verdict and Deployment

    Decision

    Whether the number you are looking at means what it appears to mean, and what you do with it.

    Builds

    A trial count for everything tried across the whole course and an explicit verdict against a promotion standard set in advance; a written handoff plan listing what must be verified before capital is committed; and the final comparison - two runs of the same pipeline on the same data, side by side: the non-ML baseline fixed in Unit 2.5 and the finished pipeline. That comparison is allowed to be unflattering. What it always shows is which parts of the workflow moved the result and which did not.

    • Multiple Testing and the Promotion Standard
    • Porting a Strategy to Live
    • The Final Comparison