Monitor usage, queries and workloads across hundreds of Microsoft Fabric Warehouses and
SQL analytics endpoints. A scheduled Fabric notebook auto-discovers every workspace/item you can
access, runs a configurable set of queries against each item's queryinsights
views, and lands the
results together in a central Lakehouse for Power BI reporting.
┌──────────────────────────────────────────────┐
│ QueryInsightsCollector notebook (scheduled) │
│ runs as the notebook identity │
└───────────────┬──────────────────────────────┘
Fabric REST API │ discover pyodbc + AAD token
list workspaces / warehouses │ ───────────────────▶ queryinsights.* views
/ lakehouse SQL endpoints │ on each Warehouse / SQL endpoint
▼
┌──────────────────────────────────────────────┐
│ Central Lakehouse (Delta tables, qi_*) │
└───────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────┐
│ Direct Lake semantic model + Power BI report │
└──────────────────────────────────────────────┘
| Path | Purpose |
|---|---|
src/query_insights_collector.py |
Single source of truth for the collector logic (cell-marked). Also importable for local unit testing. |
notebooks/QueryInsightsCollector.ipynb |
The Fabric notebook, generated from the source. Import this into Fabric. |
config/queries.json |
Flexible, user-authored queries run against each endpoint. Edit freely. |
config/settings.example.json |
Optional settings reference (values map to notebook parameters). |
powerbi/WarehouseMonitor.pbip |
Ready-to-open Power BI Project: Direct Lake semantic model + 3 sample report pages. |
powerbi/README.md |
Semantic model + DAX measures + report pages, and how to set the connection. |
tests/test_collector.py |
Unit tests for the pure logic (config parsing, predicate/watermark building, filters). |
tools/build_notebook.py |
Regenerates the .ipynb from the source. |
tools/build_powerbi.py |
Regenerates the PBIP (TMDL model + PBIR report). |
tools/collected_schemas.json |
Committed snapshot of the real Delta schemas for the auto-modelled DMV/catalog tables. |
webapp/ |
React single-page app that reads the semantic model via the Power BI executeQueries DAX API, with searchable Workspace/Database filters. |
- Discover — calls the Fabric REST API to list all accessible workspaces, then the
Warehouses (
/warehouses) and Lakehouse SQL analytics endpoints (/lakehouses→sqlEndpointProperties) in each. Allow/deny lists let you scope the fleet. - Collect — for every enabled query in
config/queries.jsonwhosetargetsmatch the item type, it connects withpyodbc(ODBC Driver 18) using an AAD access token and runs the SQL inside the item'squeryinsightsschema. - Store — each row is tagged with source metadata (
source_workspace_*,source_item_*,collected_at) and written to a Delta table:incrementalqueries MERGE on a primary key and advance a per-item watermark (with a configurable overlap window to catch queries that appear late — QI can lag ~15 min).snapshotqueries append a point-in-time capture (ideal for the aggregated views).
- Log — every (endpoint, query) outcome is written to
qi_collection_runsfor observability.
Create (or pick) a Lakehouse to hold the collected data, e.g. WarehouseMonitorLH.
In your Fabric workspace: New → Import notebook → notebooks/QueryInsightsCollector.ipynb.
Open it and attach the Lakehouse from step 1 as the default Lakehouse (Explorer pane → Add).
Upload config/queries.json to the Lakehouse Files area under Files/config/queries.json.
The notebook reads it at run time, so you can change queries without editing code. If the file is
absent, a built-in default (exec_requests_history) is used.
The first code cell is a parameters cell. Defaults work when a default Lakehouse is attached.
Override as needed — e.g. set lakehouse_abfss_path to write to a specific Lakehouse, adjust
first_run_days_back, or use workspace_allow_list / item_deny_list to scope the fleet.
Run all cells once to validate discovery and collection, then check the qi_* Delta tables and
qi_collection_runs.
The running identity needs, per warehouse/endpoint, at least Viewer to run queries and
Contributor or higher to see full query text (command / label) — this is a
queryinsights requirement. Items where the identity lacks access are simply skipped/logged.
See docs/scheduling.md. Quickest path: the notebook's built-in
Run → Schedule (e.g. hourly). For orchestration/retries/alerts, wrap it in a Data Factory
pipeline. QI retains ~30 days of history, so run at least daily.
webapp/ is a React SPA that reads the semantic model live through the Power BI
executeQueries DAX API (no gateway). It shows fleet KPIs and top-N tables with searchable
Workspace and Database filters that apply to every card and table — the same "set once,
applies everywhere" behaviour as the report's synced slicers. See
webapp/README.md for the app-registration and run steps.
Add an object to config/queries.json. Fields:
| Field | Meaning |
|---|---|
name |
Unique id (used in logs & watermarks). |
enabled |
true/false. |
targets |
Any of "warehouse", "sql_endpoint". |
destination_table |
Delta table name (qi_* recommended). |
load_mode |
"incremental" (MERGE + watermark) or "snapshot" (append point-in-time). |
incremental_column |
Datetime column driving the watermark (incremental only). |
primary_key |
Columns to de-duplicate on (incremental only). Include source_item_id. |
sql |
T-SQL against the queryinsights schema. Put {predicate} in the WHERE clause; the collector expands it to the watermark filter (incremental) or 1 = 1 (snapshot). |
continue_on_error |
If true, a failure of this query on one endpoint never aborts the run. |
python tests\test_collector.py # run unit tests (no Fabric/Spark needed)
python tools\build_notebook.py # regenerate the .ipynb after editing the source
python tools\build_powerbi.py # regenerate the PBIP (semantic model + report)The pure functions (config parsing, predicate/watermark logic, discovery filters) are unit-tested locally; Spark/pyodbc/Fabric REST calls only execute inside the Fabric runtime.