# Data Fetcher Design ## Initial Goal Create a Python module/tool that fetches daily candlestick data from the Alpaca Market Data API for a specific ticker and date range. The intended example workflow is: - ticker: `SPY`; - date range: `2026-06-01` to `2026-06-30`; - bar size: one trading day; - fields: open, high, low, close, volume; - output: one Parquet file named for the ticker, such as `SPY.parquet`. ## Current Implementation The current implementation is intentionally small: - ticker passed as a required command line argument; - end date passed with `--end-date YYYYMMDD`, defaulting to yesterday; - end date can also be derived with `--end-date-from-parquet`; - duration passed with `--duration`, defaulting to `1 W`; - uses the Alpaca Market Data API through `alpaca-py`; - reads credentials from `--api-key` / `--secret-key`, `ALPACA_API_KEY` / `ALPACA_SECRET_KEY`, or `.env`; - writes candles to a symbol-named Parquet file; - appends to an existing symbol file and keeps one row per date. Run it with: ```sh mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY ``` To override the requested range: ```sh mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY --end-date 20250605 --duration "1 M" ``` To fetch backward from the oldest date already stored in the symbol file: ```sh mise exec -- uv run python src/trading_bot/data/fetch_alpaca_daily.py SPY --end-date-from-parquet ``` `--end-date` and `--end-date-from-parquet` cannot be used together. If the symbol Parquet file does not exist or has no rows, `--end-date-from-parquet` uses today's US/Eastern date. This expects Alpaca API credentials to be available via CLI arguments, environment variables, or `.env`. By default, output is written to `data/alpaca/daily/SPY.parquet`. Use `--output-dir` to choose another directory. Manual test instructions are in [manual-test/README.md](manual-test/README.md). ## Intended Future Behavior Later, this tool should broaden configuration around data source and normalization choices. Decided storage behavior: - Parquet files are partitioned by ticker, not by date. - Each ticker should have its own Parquet file, such as `SPY.parquet`. - The trading date is used as the row key for merges. - Re-fetching a date replaces the existing row for that date in that symbol's file. Candidate output schema: | Column | Type | Description | | --- | --- | --- | | `date` | date index | Trading session date | | `symbol` | string | Asset ticker | | `open` | float | Daily open price | | `high` | float | Daily high price | | `low` | float | Daily low price | | `close` | float | Daily close price | | `volume` | integer | Daily traded volume | Open decisions: - how to handle adjusted versus unadjusted prices; - how to handle missing sessions, market holidays, and Alpaca API rate limits; - whether to add explicit feed selection, adjustment settings, or data entitlement checks.