A CLI tool that reads your dbt manifest, queries GCP BigQuery audit logs for query access patterns, and suggests optimal clustering keys for your BigQuery tables.
- Parses your
manifest.jsonto discover all materialized dbt models and sources - Queries
INFORMATION_SCHEMA.JOBS_BY_PROJECT(or a log sink dataset) to fetch recent query history - Uses
sqlglotto parse each query and extract which columns appear inWHERE,JOIN,GROUP BY, andORDER BYclauses - Scores each column per table using weighted frequency counts (WHERE/JOIN rank highest)
- Adjusts scores based on BigQuery clustering suitability of the column's data type
- Outputs up to 4 clustering key suggestions per table, ordered by score
Requires Python 3.11+.
pip install -e .For development (includes ruff, pytest, mypy):
pip install -e ".[dev]"
pre-commit installUses Application Default Credentials. Run once before using quarry:
gcloud auth application-default loginThe BigQuery service account needs bigquery.jobs.listAll permission on the project (or roles/bigquery.resourceViewer).
Analyze audit logs and generate clustering suggestions.
quarry analyze \
--manifest ./target/manifest.json \
--project my-gcp-project \
--region us \
--lookback-days 90 \
--output quarry_suggestions.yamlOptions:
| Option | Default | Description |
|---|---|---|
--manifest |
required | Path to dbt manifest.json |
--project |
required | GCP project ID |
--region |
eu |
BigQuery region (e.g. us, eu, us-central1) |
--lookback-days |
90 |
Days of query history to analyze |
--output |
quarry_suggestions.yaml |
Output file path |
--format |
yaml |
Output format: yaml or json |
--min-queries |
10 |
Minimum query count to include a table |
--query-limit |
10000 |
Max queries to fetch from audit log |
--min-gb-billed |
0 (disabled) |
Minimum GB billed to include a query (e.g. 0.5 for 500 MB) |
--config |
— | Path to quarry.toml config file |
--verbose / -v |
false |
Print debug info |
Example output (quarry_suggestions.yaml):
my-project.my_dataset.orders:
suggested_keys: [customer_id, status, created_at]
scores:
customer_id: 142.5
status: 88.0
created_at: 61.5
query_count_analyzed: 847
reasoning: "Analyzed 847 queries. `customer_id` used in WHERE: 91, JOIN: 47. ..."
original_file_path: models/orders.sqlApply clustering suggestions to dbt model SQL files by injecting {{ config(cluster_by=[...]) }} blocks.
# Preview changes (dry run, default)
quarry apply --suggestions quarry_suggestions.yaml --dbt-project .
# Write changes to disk
quarry apply --suggestions quarry_suggestions.yaml --dbt-project . --no-dry-run
# Apply only one table
quarry apply --suggestions quarry_suggestions.yaml --table my-project.my_dataset.orders --no-dry-runFor each column seen across all analyzed queries, the score is computed in two steps.
Step 1 — weighted frequency
Count how many times the column appears in each clause type across all queries, then multiply by the clause weight and sum:
score = (WHERE_count × 3.0) + (JOIN_count × 2.5) + (GROUP_BY_count × 1.5) + (ORDER_BY_count × 0.5)
The weights reflect how much each clause benefits from clustering. WHERE and JOIN filters directly prune data scanned; GROUP BY benefits less; ORDER BY barely at all.
Step 2 — data type multiplier
If the column appears in the dbt manifest with a known data type, the score is multiplied by a cardinality factor:
| Type | Multiplier | Reason |
|---|---|---|
DATE |
×1.3 | High clustering benefit |
BOOL, DATETIME |
×1.2 | High clustering benefit |
TIMESTAMP, INT64 |
×1.1 | Good clustering candidate |
NUMERIC |
×0.7 | Moderate |
FLOAT64 |
×0.5 | Too many distinct values |
BYTES |
×0.3 | Poor clustering candidate |
If the column isn't in the manifest (no type info), the multiplier defaults to ×1.0.
The top 4 columns by final score become the suggested clustering keys. Columns in the exclusion list (_partitiontime, _partitiondate, etc.) are always skipped. Weights and multipliers are fully configurable via quarry.toml.
quarry checks whether referenced tables are partitioned and whether queries are filtering on the partition column. The analyze output includes a partition_filter_rate per table. Tables with low usage (below min_partition_filter_rate, default 50%) are flagged in the CLI summary and skipped by apply — fix partition filters before clustering.
The partition column is automatically excluded from clustering key suggestions.
Configure the threshold in quarry.toml:
min_partition_filter_rate = 0.5Some tables may have a low partition filter rate for legitimate reasons (e.g. GDPR clearing jobs that must do full scans). Exclude them from the partition check with:
[excluded_tables]
partition_check = [
"my-project.my_dataset.gdpr_table",
"my-project.my_dataset.another_table",
]Excluded tables are still analyzed and clustered — only the partition filter warning and apply skip are suppressed.
information_schema (default): Queries INFORMATION_SCHEMA.JOBS_BY_PROJECT directly — no setup required beyond IAM permissions. Available in all regions.
log_sink: Queries a BigQuery dataset that receives exported Cloud Audit Logs. Requires a log sink to be configured. Useful if you want to analyze logs across projects or need longer retention.
# Run tests
pytest
# Lint + format
ruff check . --fix
ruff format .