This repository estimates daily official Fama-French factor values for dates where market data is available but the Kenneth French daily factor files have not yet been released. The active target set is the five Fama-French factors plus the daily Kenneth French momentum factor:
Mkt-RF, SMB, HML, RMW, CMA, Mom
The intended workflow is an after-close nowcast:
estimate FF5(t) using same-day market data through t
Historical official FF5 values are used as supervised labels during training. They are not used as input lag features in the default configuration. The default estimator is a market-only all-candidates ElasticNet model, selected from the release-gap backtests in this repo.
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"Run the full current workflow:
./start.shThis runs:
python -m ff5_predictor.cli nowcast --config config/nowcast/latest.yaml
python -m ff5_predictor.cli backtest-nowcast --config config/nowcast/backtest_release_gap.yaml
python -m ff5_predictor.cli build-nowcast-dataset --config config/nowcast/diagnostic.yaml
python -m ff5_predictor.cli list-modelsUse start_nowcast.sh if you want the same workflow without the final model-list print.
Latest unreleased-date estimates:
data/nowcasts/latest/latest/predictions/latest_nowcast.csv
data/nowcasts/latest/latest/predictions/official_plus_nowcast_series.csv
Latest model attribution:
data/nowcasts/latest/latest/attribution/
Release-gap backtest outputs:
data/nowcasts/market_only_elasticnet_backtest_v1/<timestamp>/predictions/release_gap_predictions.csv
data/nowcasts/market_only_elasticnet_backtest_v1/<timestamp>/metrics/model_ranking.csv
data/nowcasts/market_only_elasticnet_backtest_v1/<timestamp>/metrics/shared_date_metrics.json
Fully model-implied historical series:
data/nowcasts/model_implied_series_v1/<timestamp>/predictions/model_implied_ff5_series.csv
data/nowcasts/model_implied_series_v1/<timestamp>/predictions/model_implied_factor_series.csv
data/nowcasts/model_implied_series_v1/<timestamp>/predictions/official_minus_model_implied_series.csv
data/nowcasts/model_implied_series_v1/<timestamp>/metrics/error_summary.csv
data/nowcasts/model_implied_series_v1/<timestamp>/metrics/model_ranking.csv
Residual analysis outputs:
<run_dir>/analysis/residuals/tables/residual_panel.csv
<run_dir>/analysis/residuals/tables/residual_market_lead_lag_correlation.csv
<run_dir>/analysis/residuals/tables/residual_regime_summary.csv
<run_dir>/analysis/residuals/figures/
Diagnostic train/inference datasets:
data/nowcasts/market_only_diagnostic_v1/<timestamp>/datasets/
The default config is:
config/nowcast/latest.yaml
It uses:
- model:
elasticnet - training window: 2520 aligned rows, roughly 10 years
- inputs: same-day market/ETF OHLC-derived returns, rolling features, and proxy spreads
- labels: official factor values in decimal return units
- excluded inputs: FF5/RF lag features and recursive factor lags
- attribution: linear feature contribution artifacts for the ElasticNet model
The release-gap backtest config is:
config/nowcast/backtest_release_gap.yaml
It simulates official-data release gaps by hiding future official FF5 rows after each cutoff, fitting only on rows available at the cutoff, and predicting hidden dates with same-day market data.
To estimate a specific currently unreleased date or date range:
python -m ff5_predictor.cli nowcast \
--config config/nowcast/latest.yaml \
--start-date 2026-07-01 \
--end-date 2026-07-01To generate a historical release-gap prediction series:
python -m ff5_predictor.cli backtest-nowcast \
--config config/nowcast/backtest_release_gap.yaml \
--start-date 2020-01-01 \
--end-date 2026-03-31The historical training window is not truncated by these flags. They filter prediction target dates only.
The latest-date nowcast fills only the official release gap. To create a consistent historical series where every value is generated by the model, run:
python -m ff5_predictor.cli model-implied-series \
--config config/nowcast/model_implied_series.yamlor:
./start_model_implied_series.shThis walk-forward series trains only on dates before each checkpoint, predicts each historical target date, and saves both the model estimate and the official value for comparison. The default config predicts every aligned historical date after the minimum training period while refitting every 21 rows. Set model_implied_series.refit_step_rows: 1 for exact daily refits.
To analyze differences between official values and model-implied estimates after a run has completed:
python -m ff5_predictor.cli analyze-residuals \
--config config/nowcast/model_implied_series.yaml \
--run-dir data/nowcasts/model_implied_series_v1/<timestamp> \
--model-type elasticnetThis writes residual summary tables, cross-factor residual correlations, residual autocorrelations, market lead/lag correlations, regime summaries, simple market regressions, and SVG figures. The standard market overlay figure pairs SPY cumulative return with mean and max absolute residuals in basis points. For release-gap backtests, use --release-gap-size 1 or --gap-day 1 if you want a single comparable gap horizon.
Fama-French data is loaded from getFamaFrenchFactors when practical, otherwise from the official Kenneth French remote zip. Existing repository-local FF5 CSV or zip files are ignored.
Market data is downloaded from yfinance with auto_adjust=True and cached as Parquet. If a requested nowcast date is beyond cached market coverage, the loader attempts a fresh yfinance download before deciding no prediction is possible.
All factor values are stored internally as decimal returns.
Generate a standalone browser-based FF5 viewer:
python -m ff5_predictor.cli view-factors --config config/default.yaml --open-browserOverlay a nowcast/backtest run:
python -m ff5_predictor.cli view-factors \
--config config/nowcast/backtest_release_gap.yaml \
--run-dir data/nowcasts/market_only_elasticnet_backtest_v1/latest \
--model-type elasticnet \
--gap-day 1 \
--open-browserExploratory configs and runner scripts are retained for reproducibility, but they are not the recommended workflow:
config/research/
scripts/research/
These include feature extraction, clustered Ridge, refined ElasticNet, per-factor ElasticNet, size/value universe variants, and TFT experiments.
.venv/bin/python -m pytest