LuigiCerone/quarry

A CLI tool that reads your dbt manifest, queries GCP BigQuery audit logs for query access patterns, and suggests optimal clustering keys for your BigQuery tables.

★ 0Forks 0PythonGitHub ↗Compare
bigquerycost-optimizationdbt

README

quarry

A CLI tool that reads your dbt manifest, queries GCP BigQuery audit logs for query access patterns, and suggests optimal clustering keys for your BigQuery tables.

How it works

  1. Parses your manifest.json to discover all materialized dbt models and sources
  2. Queries INFORMATION_SCHEMA.JOBS_BY_PROJECT (or a log sink dataset) to fetch recent query history
  3. Uses sqlglot to parse each query and extract which columns appear in WHERE, JOIN, GROUP BY, and ORDER BY clauses
  4. Scores each column per table using weighted frequency counts (WHERE/JOIN rank highest)
  5. Adjusts scores based on BigQuery clustering suitability of the column's data type
  6. Outputs up to 4 clustering key suggestions per table, ordered by score

Installation

Requires Python 3.11+.

pip install -e .

For development (includes ruff, pytest, mypy):

pip install -e ".[dev]"
pre-commit install

GCP Authentication

Uses Application Default Credentials. Run once before using quarry:

gcloud auth application-default login

The BigQuery service account needs bigquery.jobs.listAll permission on the project (or roles/bigquery.resourceViewer).

Commands

quarry analyze

Analyze audit logs and generate clustering suggestions.

quarry analyze \
  --manifest ./target/manifest.json \
  --project my-gcp-project \
  --region us \
  --lookback-days 90 \
  --output quarry_suggestions.yaml

Options:

Option Default Description
--manifest required Path to dbt manifest.json
--project required GCP project ID
--region eu BigQuery region (e.g. us, eu, us-central1)
--lookback-days 90 Days of query history to analyze
--output quarry_suggestions.yaml Output file path
--format yaml Output format: yaml or json
--min-queries 10 Minimum query count to include a table
--query-limit 10000 Max queries to fetch from audit log
--min-gb-billed 0 (disabled) Minimum GB billed to include a query (e.g. 0.5 for 500 MB)
--config — Path to quarry.toml config file
--verbose / -v false Print debug info

Example output (quarry_suggestions.yaml):

my-project.my_dataset.orders:
  suggested_keys: [customer_id, status, created_at]
  scores:
    customer_id: 142.5
    status: 88.0
    created_at: 61.5
  query_count_analyzed: 847
  reasoning: "Analyzed 847 queries. `customer_id` used in WHERE: 91, JOIN: 47. ..."
  original_file_path: models/orders.sql

quarry apply

Apply clustering suggestions to dbt model SQL files by injecting {{ config(cluster_by=[...]) }} blocks.

# Preview changes (dry run, default)
quarry apply --suggestions quarry_suggestions.yaml --dbt-project .

# Write changes to disk
quarry apply --suggestions quarry_suggestions.yaml --dbt-project . --no-dry-run

# Apply only one table
quarry apply --suggestions quarry_suggestions.yaml --table my-project.my_dataset.orders --no-dry-run

How scoring works

For each column seen across all analyzed queries, the score is computed in two steps.

Step 1 — weighted frequency

Count how many times the column appears in each clause type across all queries, then multiply by the clause weight and sum:

score = (WHERE_count × 3.0) + (JOIN_count × 2.5) + (GROUP_BY_count × 1.5) + (ORDER_BY_count × 0.5)

The weights reflect how much each clause benefits from clustering. WHERE and JOIN filters directly prune data scanned; GROUP BY benefits less; ORDER BY barely at all.

Step 2 — data type multiplier

If the column appears in the dbt manifest with a known data type, the score is multiplied by a cardinality factor:

Type Multiplier Reason
DATE ×1.3 High clustering benefit
BOOL, DATETIME ×1.2 High clustering benefit
TIMESTAMP, INT64 ×1.1 Good clustering candidate
NUMERIC ×0.7 Moderate
FLOAT64 ×0.5 Too many distinct values
BYTES ×0.3 Poor clustering candidate

If the column isn't in the manifest (no type info), the multiplier defaults to ×1.0.

The top 4 columns by final score become the suggested clustering keys. Columns in the exclusion list (_partitiontime, _partitiondate, etc.) are always skipped. Weights and multipliers are fully configurable via quarry.toml.

Partition filter check

quarry checks whether referenced tables are partitioned and whether queries are filtering on the partition column. The analyze output includes a partition_filter_rate per table. Tables with low usage (below min_partition_filter_rate, default 50%) are flagged in the CLI summary and skipped by apply — fix partition filters before clustering.

The partition column is automatically excluded from clustering key suggestions.

Configure the threshold in quarry.toml:

min_partition_filter_rate = 0.5

Some tables may have a low partition filter rate for legitimate reasons (e.g. GDPR clearing jobs that must do full scans). Exclude them from the partition check with:

[excluded_tables]
partition_check = [
    "my-project.my_dataset.gdpr_table",
    "my-project.my_dataset.another_table",
]

Excluded tables are still analyzed and clustered — only the partition filter warning and apply skip are suppressed.

Audit log sources

information_schema (default): Queries INFORMATION_SCHEMA.JOBS_BY_PROJECT directly — no setup required beyond IAM permissions. Available in all regions.

log_sink: Queries a BigQuery dataset that receives exported Cloud Audit Logs. Requires a log sink to be configured. Useful if you want to analyze logs across projects or need longer retention.

Development

# Run tests
pytest

# Lint + format
ruff check . --fix
ruff format .

Contributors

LuigiCerone

Issues