- Python 3.12
- Internet connection (SEC filings are downloaded on demand).
- CSV inputs:
BS_Q.csv(filings) andsub_map.csv(CIK-to-submission mapping). Sample files live underdata/samples/.- Use
data/samples/BS_Q.sample.csvanddata/samples/sub_map.sample.csvas templates to confirm the expected columns before dropping in your own data.
- Use
python -m venv .venv
source .venv/bin/activate # On Windows use: .venv\Scripts\activate
pip install pandas requests beautifulsoup4 tqdm sec-apiAdd your SEC Extractor API key to a .env file in the project root:
echo "SEC_API_KEY=your_sec_api_key_here" >> .envYou can also copy .env.sample to .env and replace the placeholder value; keep the .env file at the project root so scripts can load the key automatically.
The SEC_API_KEY is required to retrieve Item 1A word counts via the SEC API.
src/ Core pipeline modules and shared helpers.
scripts/ CLI utilities (enrichment, data quality tools, etc.).
data/ Workspace for inputs/outputs (ignored by git).
Run the pipeline entry point with your filings CSV plus the submission map. The command below produces data/outputs/bsq_quarter.final.csv together with intermediate checkpoints.
python3 main.py --bsq data/samples/BS_Q.csv --submap data/samples/sub_map.csvKey flags:
--bsq– input filings CSV (requires columns such ascik,filingUrl,filedAt).--submap– lookup table used to decide the fiscal quarter.--max-rows– optional cap for smoke tests.
Feed the quarter CSV into the Item 1A enricher. It drops rows that lack a ticker, fetches the risk section via the SEC Extractor API (same source for keywords and word counts), computes totals, deduplicates repeated filings, sorts the final CSV alphabetically by ticker (when present), and issues SEC requests concurrently whenever --rate 0 (the default) to minimize wall-clock time.
python3 scripts/enrich_item1a.py --input data/outputs/bsq_quarter.final.csvUseful options:
--output– override the destination CSV (defaultdata/outputs/bsq_quarter.item1a.csv).--rate– seconds to sleep between requests; set> 0to force sequential throttling.--max-workers– cap the number of parallel SEC requests when--rate 0(default auto, max 8).--max-rows– run a quick check against the first N filings.--keep-text– retain the raw Item 1A text in the output.--no-dedupe– skip(cik, fyear, quarter)deduplication if you want every row.
python3 scripts/enrich_item1a.py --input data/samples/BS_Q_test.csv --max-rows 5