Files
trading-bot/docs/training-dataset.md
T
2026-07-26 21:38:14 +03:00

4.7 KiB

Training Dataset Design

Initial Goal

Build a supervised learning dataset from stored IBKR daily candle data. The first dataset predicts whether SPY closes higher five trading days after the observation date.

The first implementation is ../notebooks/spy_direction_dataset.ipynb, a Python notebook using pandas. This keeps the feature calculations inspectable while the dataset design is still changing. Once the feature contract settles, reusable loading and feature-building code can move into src/trading_bot/models or src/trading_bot/data.

Raw Data

Raw daily candles are stored under:

data/ibkr/daily

Storage expectations:

  • one Parquet file per symbol;
  • file name is the symbol, such as SPY.parquet;
  • each row contains one daily candle for that symbol;
  • expected columns are date, symbol, open, high, low, close, and volume.

The initial dataset requires at least these symbols:

  • SPY;
  • VIX;
  • TLT;
  • USO.

Observation Rows

Each dataset row represents one trading date. Feature values are derived from data available on or before that date. The target is derived from SPY closing price five trading days after that date.

Rows should be dropped when any required feature or target value cannot be computed.

Input Features

Feature Description
SPY_ret_5 SPY fractional return over the previous 5 trading days.
SPY_ret_20 SPY fractional return over the previous 20 trading days.
SPY_dist_sma50 SPY distance from its 50 trading day simple moving average, divided by that average.
VIX_change_5 VIX fractional change over the previous 5 trading days.
VIX_rank_20 VIX percentile rank within a rolling 20 trading day window.
TLT_ret_10 TLT fractional return over the previous 10 trading days.
USO_ret_5 USO fractional return over the previous 5 trading days.
SPY_TLT_ratio_ret Fractional 5 trading day change in the SPY to TLT close-price ratio.

Fractional percentage changes should follow pandas pct_change convention. For example, a five percent return is represented as 0.05.

Output Target

The target column is a binary indicator of SPY forward return over the next five trading days.

Target Description
spy_up_5d 1.0 when SPY closes above today's close five trading days later; 0.0 when SPY closes below today's close five trading days later; 0.5 when the future close equals today's close.

Using 0.5 for unchanged prices preserves the row while making the target explicitly neutral.

Data Alignment

The dataset should use trading-day alignment rather than calendar-day alignment. A five day lookback or forecast horizon means five available trading sessions for the relevant symbol, not five calendar days.

SPY should define the primary observation calendar because the first target is an SPY forward-return label. Other symbols should be joined to the SPY observation dates after their own features are computed.

Notebook Feature Sketch

The first notebook should use pandas transformations close to this shape:

import numpy as np
import pandas as pd

df["SPY_ret_5"] = df["SPY_close"].pct_change(5)
df["SPY_ret_20"] = df["SPY_close"].pct_change(20)

sma_50 = df["SPY_close"].rolling(50).mean()
df["SPY_dist_sma50"] = (df["SPY_close"] - sma_50) / sma_50

df["VIX_change_5"] = df["VIX_close"].pct_change(5)
df["VIX_rank_20"] = df["VIX_close"].rolling(20).rank(pct=True)

df["TLT_ret_10"] = df["TLT_close"].pct_change(10)
df["USO_ret_5"] = df["USO_close"].pct_change(5)
df["SPY_TLT_ratio_ret"] = (df["SPY_close"] / df["TLT_close"]).pct_change(5)

spy_forward_close = df["SPY_close"].shift(-5)
df["spy_up_5d"] = np.select(
    [
        spy_forward_close > df["SPY_close"],
        spy_forward_close < df["SPY_close"],
    ],
    [1.0, 0.0],
    default=0.5,
)

df_model = df.dropna().copy()

Dataset Splits

Training, validation, and test splits should be chronological:

  • oldest rows for training;
  • newer rows for validation;
  • newest rows for the final test set.

A split indicator column such as split is useful in the dataset artifact for auditability and reproducibility. It should be treated as metadata, not as a model input feature. The model training code should build X from the explicit feature column list and exclude metadata columns such as date, split, raw close prices, and the target.

Open Decisions

  • exact adjusted versus unadjusted close handling;
  • whether same-day VIX, TLT, and USO values are acceptable for the intended trading decision timing;
  • exact train, validation, and test date boundaries or split percentages;
  • output dataset file location, schema metadata, and versioning.