Files
trading-bot/docs/data-fetcher.md
T
2026-08-11 20:23:59 +03:00

2.9 KiB

Data Fetcher Design

Initial Goal

Create a Python module/tool that fetches daily candlestick data from the Alpaca Market Data API for a specific ticker and date range.

The intended example workflow is:

  • ticker: SPY;
  • date range: 2026-06-01 to 2026-06-30;
  • bar size: one trading day;
  • fields: open, high, low, close, volume;
  • output: one Parquet file named for the ticker, such as SPY.parquet.

Current Implementation

The current implementation is intentionally small:

  • ticker passed as a required command line argument;
  • end date passed with --end-date YYYYMMDD, defaulting to yesterday;
  • end date can also be derived with --end-date-from-parquet;
  • duration passed with --duration, defaulting to 1 W;
  • uses the Alpaca Market Data API through alpaca-py;
  • reads credentials from --api-key / --secret-key, ALPACA_API_KEY / ALPACA_SECRET_KEY, or .env;
  • writes candles to a symbol-named Parquet file;
  • appends to an existing symbol file and keeps one row per date.

Run it with:

mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY

To override the requested range:

mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY --end-date 20250605 --duration "1 M"

To fetch backward from the oldest date already stored in the symbol file:

mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY --end-date-from-parquet

--end-date and --end-date-from-parquet cannot be used together. If the symbol Parquet file does not exist or has no rows, --end-date-from-parquet uses today's US/Eastern date.

This expects Alpaca API credentials to be available via CLI arguments, environment variables, or .env.

By default, output is written to data/alpaca/daily/SPY.parquet. Use --output-dir to choose another directory.

Manual test instructions are in manual-test/README.md.

Intended Future Behavior

Later, this tool should broaden configuration around data source and normalization choices.

Decided storage behavior:

  • Parquet files are partitioned by ticker, not by date.
  • Each ticker should have its own Parquet file, such as SPY.parquet.
  • The trading date is used as the row key for merges.
  • Re-fetching a date replaces the existing row for that date in that symbol's file.

Candidate output schema:

Column Type Description
date date index Trading session date
symbol string Asset ticker
open float Daily open price
high float Daily high price
low float Daily low price
close float Daily close price
volume integer Daily traded volume

Open decisions:

  • how to handle adjusted versus unadjusted prices;
  • how to handle missing sessions, market holidays, and Alpaca API rate limits;
  • whether to add explicit feed selection, adjustment settings, or data entitlement checks.