ML4T Backtest¶
Event-driven backtesting engine with configurable execution semantics, validated against four independent frameworks.
Use ml4t-backtest when notebook research is no longer enough and you need explicit,
reproducible answers to practical execution questions: when orders fill, how stops trigger,
how cash is reserved, and how results change when you match another framework's behavior.
-
Run Your First Backtest --- Run a complete strategy on bundled synthetic bars and inspect its fills. Quickstart
-
User Guide --- Find the workflow for data, strategies, orders, account rules, costs, risk, and result analysis. User Guide
-
Validated Against 4 Frameworks --- Compare retained scenario evidence for VectorBT, Backtrader, and Zipline, plus native and Chapter 16 evidence for LEAN. Profiles
-
Chapters 16-19 --- The book develops the ideas in notebooks. This library turns them into reusable execution and reporting workflows. Book Guide
Overview¶
ml4t-backtest is the simulation layer in the ML4T stack. It sits between research and
deployment:
ml4t-dataprepares canonical market datasetsml4t-engineerproduces labels and featuresml4t-diagnosticvalidates signals, models, and portfolio behaviorml4t-backtestsimulates execution with explicit, configurable semanticsml4t-livereuses the same strategy surface for paper and live rollout
Quick Example¶
This run uses a bundled synthetic AAPL panel, so it needs no API key or data download. The first backtest walks through the orders and result records.
from importlib.metadata import version
import polars as pl
from ml4t.backtest import BacktestConfig, DataFeed, Engine
from ml4t.backtest.example_data import ExampleRoundTrip, load_example_prices
prices = load_example_prices("equity").filter(pl.col("asset") == "AAPL")
result = Engine(
DataFeed(prices_df=prices),
ExampleRoundTrip("AAPL", 100),
BacktestConfig(initial_cash=100_000),
).run()
print("ml4t-backtest " + version("ml4t-backtest"))
print(f"fills={len(result.fills)} final=${result.metrics['final_value']:.2f}")
Each Engine instance is single-use. Create a new instance for every independent run.
Moving a Zipline strategy? Follow the task-level migration map and run its checked target-weight example before comparing framework results.
The convenience function accepts the same price panel and strategy directly:
import polars as pl
from ml4t.backtest import run_backtest
from ml4t.backtest.example_data import ExampleRoundTrip, load_example_prices
prices = load_example_prices("equity").filter(pl.col("asset") == "AAPL")
result = run_backtest(prices, ExampleRoundTrip("AAPL", 100), config="backtrader")
print(f"fills={len(result.fills)} final=${result.metrics['final_value']:.2f}")
Why ML4T Backtest?¶
Configurable execution semantics. Framework differences such as fill ordering, stop modes, cash policies, and settlement are named configuration parameters. Retained evidence below states which pinned framework scenarios currently match exactly.
Quote-aware when you need it. The feed can cache bid, ask, midpoint, and quote sizes additively. Market execution and position marking can use price, bid, ask, quote_mid, or quote_side.
Retained validation evidence. The primary audit uses five real-data strategy workloads with frozen inputs. Separate synthetic scenario and 250-asset stress suites isolate conventions and exercise high event counts.
| Feature | Description |
|---|---|
| Event-driven | Explicit decision and fill timing; same-bar settings require a causal-data check |
| Configurable behavior | Set fill timing, cash, costs, and order processing explicitly |
| Quote-aware execution | Side-aware fills and separate mark pricing |
| 10 framework profiles | Configure VectorBT, Backtrader, Zipline, and LEAN semantics |
| Risk management | Stop-loss, take-profit, trailing stops, portfolio limits |
| Multi-asset | Rebalancing, weight targets, exit-first ordering |
| Rich persistence | Export trades, fills, equity, portfolio state, and daily P&L to Parquet |
Parity Validation¶
Real-strategy comparisons run frozen market data and model-derived targets through each supported engine pair. The synthetic scenario and stress suites provide narrower conformance evidence.
Real-strategy audit¶
17/17 required pairs pass; 8 pairs are declared unsupported. The audit uses five real-data strategy workloads with frozen historical market data and model-derived targets. A pass requires identical valuation timestamp coverage, complete fill streams with quantities equal at 1e-5 and prices equal at 1e-8, and account monetary values that round to the same cent. The FX workload uses the USD-quoted pairs in its frozen target stream so every required engine uses native USD valuation.
The parity protocol disables transaction costs and position rules on both sides. It tests target sizing, order sequencing, fills, cash and margin behavior, funding where applicable, and valuation. It does not claim to reproduce each selected case-study production result with its original costs and risk overlays.
| Real strategy | Pinned framework | Current result | Evidence |
|---|---|---|---|
| ETF allocation | VectorBT Pro 2026.6.27 | fills equal at declared field precision; 1,995 valuations and terminal exact at 1e-8 | real-strategy evidence |
| ETF allocation | VectorBT OSS 1.1.0 | fills equal at declared field precision; 1,995 valuations and terminal exact at 1e-8 | real-strategy evidence |
| ETF allocation | Backtrader 1.9.78.123 | fills equal at declared field precision; 1,995 valuations and terminal exact at 1e-8 | real-strategy evidence |
| ETF allocation | Zipline Reloaded 3.1.1 | fills equal at declared field precision; 1,995 valuations and terminal exact at 1e-8 | real-strategy evidence |
| ETF allocation | LEAN 18001 | fills equal at declared field precision; 1,995 valuations and terminal exact at 1e-8 | real-strategy evidence |
| CME futures | VectorBT Pro 2026.6.27 | fills equal at declared field precision; 1,595 valuations within $0.01 (max raw gap $0.00000010); terminal within $0.01 (raw gap $0.00000007) | real-strategy evidence |
| CME futures | Backtrader 1.9.78.123 | fills equal at declared field precision; 1,595 valuations within $0.01 (max raw gap $0.00000015); terminal within $0.01 (raw gap $0.00000015) | real-strategy evidence |
| Crypto perpetual funding | LEAN 18001 | fills equal at declared field precision; 2,426 valuations and terminal exact at 1e-8 | real-strategy evidence |
| FX allocation (USD-quoted pairs) | VectorBT Pro 2026.6.27 | fills equal at declared field precision; 2,108 valuations and terminal exact at 1e-8 | real-strategy evidence |
| FX allocation (USD-quoted pairs) | VectorBT OSS 1.1.0 | fills equal at declared field precision; 2,108 valuations and terminal exact at 1e-8 | real-strategy evidence |
| FX allocation (USD-quoted pairs) | Backtrader 1.9.78.123 | fills equal at declared field precision; 2,108 valuations and terminal exact at 1e-8 | real-strategy evidence |
| FX allocation (USD-quoted pairs) | LEAN 18001 | fills equal at declared field precision; 2,108 valuations and terminal exact at 1e-8 | real-strategy evidence |
| US equity panel | VectorBT Pro 2026.6.27 | fills equal at declared field precision; 4,146 valuations within $0.01 (max raw gap $0.00001950); terminal within $0.01 (raw gap $0.00001880) | real-strategy evidence |
| US equity panel | VectorBT OSS 1.1.0 | fills equal at declared field precision; 4,146 valuations within $0.01 (max raw gap $0.00001910); terminal within $0.01 (raw gap $0.00001870) | real-strategy evidence |
| US equity panel | Backtrader 1.9.78.123 | fills equal at declared field precision; 4,146 valuations within $0.01 (max raw gap $0.00000170); terminal within $0.01 (raw gap $0.00000160) | real-strategy evidence |
| US equity panel | Zipline Reloaded 3.1.1 | fills equal at declared field precision; 4,027 valuations within $0.01 (max raw gap $0.00000190); terminal within $0.01 (raw gap $0.00000030) | real-strategy evidence |
| US equity panel | LEAN 18001 | fills equal at declared field precision; 4,027 valuations within $0.01 (max raw gap $0.00000460); terminal within $0.01 (raw gap $0.00000420) | real-strategy evidence |
Real-strategy engine performance¶
The table reports engine-call wall time for all 17 correctness-passing pairs. The ratio is framework median / ML4T median; values above 1 mean ML4T completed the engine call faster.
| Real strategy | Pinned framework | Framework median (95% CI), s | ML4T median (95% CI), s | Framework / ML4T median |
|---|---|---|---|---|
| ETF allocation | VectorBT Pro 2026.6.27 | 0.293 (0.290-0.298) | 0.424 (0.421-0.434) | 0.691x |
| ETF allocation | VectorBT OSS 1.1.0 | 0.192 (0.176-0.493) | 0.414 (0.413-0.417) | 0.462x |
| ETF allocation | Backtrader 1.9.78.123 | 9.631 (9.428-9.710) | 0.439 (0.434-0.444) | 21.954x |
| ETF allocation | Zipline Reloaded 3.1.1 | 3.888 (3.878-3.959) | 0.621 (0.618-0.626) | 6.256x |
| ETF allocation | LEAN 18001 | 2.638 (2.516-2.707) | 0.750 (0.743-0.757) | 3.520x |
| CME futures | VectorBT Pro 2026.6.27 | 0.290 (0.289-0.294) | 0.410 (0.406-0.958) | 0.708x |
| CME futures | Backtrader 1.9.78.123 | 6.639 (5.831-8.332) | 0.428 (0.422-1.235) | 15.504x |
| Crypto perpetual funding | LEAN 18001 | 2.828 (2.737-2.952) | 1.327 (0.657-2.183) | 2.131x |
| FX allocation (USD-quoted pairs) | VectorBT Pro 2026.6.27 | 0.288 (0.286-0.294) | 0.148 (0.147-0.149) | 1.947x |
| FX allocation (USD-quoted pairs) | VectorBT OSS 1.1.0 | 0.144 (0.142-0.146) | 0.145 (0.145-0.146) | 0.988x |
| FX allocation (USD-quoted pairs) | Backtrader 1.9.78.123 | 0.434 (0.430-0.440) | 0.147 (0.146-0.149) | 2.947x |
| FX allocation (USD-quoted pairs) | LEAN 18001 | 0.953 (0.936-1.164) | 0.158 (0.155-0.164) | 6.033x |
| US equity panel | VectorBT Pro 2026.6.27 | 3.806 (3.336-4.341) | 23.306 (21.906-25.833) | 0.163x |
| US equity panel | VectorBT OSS 1.1.0 | 17.116 (17.069-17.633) | 25.738 (24.983-27.163) | 0.665x |
| US equity panel | Backtrader 1.9.78.123 | 668.611 (642.993-718.202) | 26.887 (26.406-29.406) | 24.868x |
| US equity panel | Zipline Reloaded 3.1.1 | 117.445 (116.308-118.391) | 26.821 (25.473-27.644) | 4.379x |
| US equity panel | LEAN 18001 | 48.436 (48.221-48.986) | 27.688 (27.575-27.796) | 1.749x |
Measured 2026-09-24 on Linux-6.8.0-139-generic-x86_64-with-glibc2.39 with 24 logical CPUs. Each side used one isolated warm-up process and ten isolated measured processes. The timer includes only the engine call; it excludes input loading, model inference, target construction, adapter preparation, result extraction, serialization, reporting. These measurements apply only to the named strategy, framework version, frozen input bundle, and machine. Raw samples and bootstrap intervals are retained in real-strategy performance evidence.
Synthetic diagnostic scenarios¶
The scenario matrix contains synthetic conformance tests. "Exact" means terminal values, ordered closed trades, and ordered fills match after 1e-8 quantization. Each record declares whether a surface is native, reconstructed, aggregate-only, input-only, or unavailable. These results test isolated conventions, not realistic strategy equivalence.
| Profile | Pinned framework | Required scenarios | Evidence |
|---|---|---|---|
vectorbt_strict |
VectorBT Pro 2026.6.27 | 17/17 exact | scenario evidence |
vectorbt_oss_strict |
VectorBT OSS 1.1.0 | 16/16 exact | scenario evidence |
backtrader_strict |
Backtrader 1.9.78.123 | 17/17 exact | scenario evidence |
zipline_strict |
Zipline Reloaded 3.1.1 | 16/16 exact | scenario evidence |
The synthetic stress workload contains 250 assets and 5,040 daily sessions (1,260,000 bars). Every row has zero canonical gap for target intents, native fills, closed trades reconstructed from those fills, and terminal state reconstructed from the fill ledger and final marks. Fill records use 1e-8 precision; monetary totals use cent precision.
| Profile | Current framework | Target intents | Native fills | Fill-derived closed trades | Terminal value | Evidence |
|---|---|---|---|---|---|---|
vectorbt_strict |
VectorBT Pro 2026.6.27 | 427,790 | 423,313 | 222,751 | 1,285,886.320000 | scale evidence |
vectorbt_oss_strict |
VectorBT OSS 1.1.0 | 427,790 | 417,941 | 211,322 | 1,345,348.850000 | scale evidence |
backtrader_strict |
Backtrader 1.9.78.123 | 427,790 | 343,813 | 182,019 | -9,166,273.560000 | scale evidence |
zipline_strict |
Zipline Reloaded 3.1.1 | 427,790 | 427,696 | 226,434 | 10,504,095.900000 | scale evidence |
lean |
LEAN 18001 | 427,790 | 361,297 | 191,297 | 184,538.130000 | scale evidence |
Release performance evidence runs deterministic single-asset, 250-asset daily, quote-aware, rebalance, and partial-fill workloads in isolated processes. It reports setup and engine runtime separately, measures whole-process peak RSS, and verifies retained financial-output checksums and counts. The 250-asset workload periodically enters and exits 50 positions. Sample spread is reported for diagnosis, while the instrument-free hotpath benchmark enforces the runtime regression limit. These ML4T-only measurements are regression baselines, not cross-machine performance claims.
The separate cross-framework performance artifact retains ten isolated measurements per runner after one warm-up. It reports complete-process wall time, process-tree and LEAN-container peak RSS, raw samples, 95% bootstrap intervals, exact output checksums, and semantic disclosures. It is retained as supporting audit evidence. The published table above instead reports engine-only timings for the realistic, correctness-passing workloads and states the workload, versions, host, date, and uncertainty needed to interpret each ratio.
Installation¶
Next Steps¶
- New here? Start with the Quickstart
- Coming from the book? Use the Book Guide
- Debugging execution behavior? Read Execution Semantics
- Need exact interfaces? See the API Reference
From Book to Library¶
If you are reading Machine Learning for Trading, Third Edition, use the docs in this order:
- learn the execution or reporting concept in the notebook
- use the Book Guide to find the matching production workflow
- move to the relevant user-guide page for the reusable API
- finish in the API Reference for exact call signatures
This is especially important for quote-aware execution, realistic reporting, rebalancing,
and strategy portability into ml4t-live.
Part of the ML4T Ecosystem¶
The same Strategy class works in both backtest and live trading via ml4t-live.