Training data generation.

This commit is contained in:
2026-07-26 21:38:14 +03:00
parent e2eb5dd574
commit 7943f48d47
7 changed files with 1389 additions and 4 deletions
+2 -1
View File
@@ -12,6 +12,8 @@ Responsible for acquiring and storing market, asset, and any future feature data
The first planned tool is an IBKR daily candle fetcher. It should fetch open, high, low, close, and volume data for a ticker and date range, then eventually persist that data to a ticker-named Parquet file. See [data-fetcher.md](data-fetcher.md).
The initial supervised training dataset is documented in [training-dataset.md](training-dataset.md). It derives market-regime features from `SPY`, `VIX`, `TLT`, and `USO`, then labels whether `SPY` closes higher five trading days later.
Parquet files are partitioned by ticker, not by date.
Open decisions:
@@ -30,7 +32,6 @@ Responsible for building datasets, training models, evaluating candidates, and w
Open decisions:
- prediction target;
- feature set;
- model family;
- validation strategy;
- evaluation metrics;
+117
View File
@@ -0,0 +1,117 @@
# Training Dataset Design
## Initial Goal
Build a supervised learning dataset from stored IBKR daily candle data. The first dataset predicts whether `SPY` closes higher five trading days after the observation date.
The first implementation is [../notebooks/spy_direction_dataset.ipynb](../notebooks/spy_direction_dataset.ipynb), a Python notebook using pandas. This keeps the feature calculations inspectable while the dataset design is still changing. Once the feature contract settles, reusable loading and feature-building code can move into `src/trading_bot/models` or `src/trading_bot/data`.
## Raw Data
Raw daily candles are stored under:
```text
data/ibkr/daily
```
Storage expectations:
- one Parquet file per symbol;
- file name is the symbol, such as `SPY.parquet`;
- each row contains one daily candle for that symbol;
- expected columns are `date`, `symbol`, `open`, `high`, `low`, `close`, and `volume`.
The initial dataset requires at least these symbols:
- `SPY`;
- `VIX`;
- `TLT`;
- `USO`.
## Observation Rows
Each dataset row represents one trading date. Feature values are derived from data available on or before that date. The target is derived from `SPY` closing price five trading days after that date.
Rows should be dropped when any required feature or target value cannot be computed.
## Input Features
| Feature | Description |
| --- | --- |
| `SPY_ret_5` | `SPY` fractional return over the previous 5 trading days. |
| `SPY_ret_20` | `SPY` fractional return over the previous 20 trading days. |
| `SPY_dist_sma50` | `SPY` distance from its 50 trading day simple moving average, divided by that average. |
| `VIX_change_5` | `VIX` fractional change over the previous 5 trading days. |
| `VIX_rank_20` | `VIX` percentile rank within a rolling 20 trading day window. |
| `TLT_ret_10` | `TLT` fractional return over the previous 10 trading days. |
| `USO_ret_5` | `USO` fractional return over the previous 5 trading days. |
| `SPY_TLT_ratio_ret` | Fractional 5 trading day change in the `SPY` to `TLT` close-price ratio. |
Fractional percentage changes should follow pandas `pct_change` convention. For example, a five percent return is represented as `0.05`.
## Output Target
The target column is a binary indicator of `SPY` forward return over the next five trading days.
| Target | Description |
| --- | --- |
| `spy_up_5d` | `1.0` when `SPY` closes above today's close five trading days later; `0.0` when `SPY` closes below today's close five trading days later; `0.5` when the future close equals today's close. |
Using `0.5` for unchanged prices preserves the row while making the target explicitly neutral.
## Data Alignment
The dataset should use trading-day alignment rather than calendar-day alignment. A five day lookback or forecast horizon means five available trading sessions for the relevant symbol, not five calendar days.
`SPY` should define the primary observation calendar because the first target is an `SPY` forward-return label. Other symbols should be joined to the `SPY` observation dates after their own features are computed.
## Notebook Feature Sketch
The first notebook should use pandas transformations close to this shape:
```python
import numpy as np
import pandas as pd
df["SPY_ret_5"] = df["SPY_close"].pct_change(5)
df["SPY_ret_20"] = df["SPY_close"].pct_change(20)
sma_50 = df["SPY_close"].rolling(50).mean()
df["SPY_dist_sma50"] = (df["SPY_close"] - sma_50) / sma_50
df["VIX_change_5"] = df["VIX_close"].pct_change(5)
df["VIX_rank_20"] = df["VIX_close"].rolling(20).rank(pct=True)
df["TLT_ret_10"] = df["TLT_close"].pct_change(10)
df["USO_ret_5"] = df["USO_close"].pct_change(5)
df["SPY_TLT_ratio_ret"] = (df["SPY_close"] / df["TLT_close"]).pct_change(5)
spy_forward_close = df["SPY_close"].shift(-5)
df["spy_up_5d"] = np.select(
[
spy_forward_close > df["SPY_close"],
spy_forward_close < df["SPY_close"],
],
[1.0, 0.0],
default=0.5,
)
df_model = df.dropna().copy()
```
## Dataset Splits
Training, validation, and test splits should be chronological:
- oldest rows for training;
- newer rows for validation;
- newest rows for the final test set.
A split indicator column such as `split` is useful in the dataset artifact for auditability and reproducibility. It should be treated as metadata, not as a model input feature. The model training code should build `X` from the explicit feature column list and exclude metadata columns such as `date`, `split`, raw close prices, and the target.
## Open Decisions
- exact adjusted versus unadjusted close handling;
- whether same-day `VIX`, `TLT`, and `USO` values are acceptable for the intended trading decision timing;
- exact train, validation, and test date boundaries or split percentages;
- output dataset file location, schema metadata, and versioning.