Quality (dw-quality)
The Quality agent monitors your data across five dimensions — completeness, accuracy, consistency, freshness, and uniqueness — and rolls them into a weighted 0–100 score per dataset, with the breakdown and trend visible so you know why a score moved. Anomaly detection is statistical (z-score against 14-day baselines), and detected anomalies are deduplicated hard: 50–100 raw signals typically collapse to 5–10 actionable ones, because an alert feed nobody reads is worse than no alert feed.
Its operating discipline matters as much as its math. It re-verifies a failure before alerting on it, distinguishes “the data is wrong” from “the check is wrong”, and never silently corrects, filters, or imputes data to make a check pass — any correction goes through the governed write path with human approval. Genuine breaks are handed to the Incidents agent rather than absorbed into monitoring.
Key capabilities
Section titled “Key capabilities”- Full quality profiling.
run_quality_checkprofiles null rates, uniqueness, distributions, referential integrity, freshness, and volume for a dataset — all columns or a chosen subset. - A score you can interrogate.
get_quality_scorereturns the 0–100 score with its per-dimension breakdown and trend, not just a number. - Deduplicated anomaly detection.
get_anomalieslists detected anomalies classified by severity (critical / warning / info), deduplicated to the actionable set by default. - Quality SLAs.
set_sladefines metric thresholds with severity levels per dataset; violations are designed to trigger alerts within a minute. - Tests born with the pipeline.
create_quality_tests_for_pipelinegenerates quality test specs for a pipeline — this is what the Pipelines agent calls so new pipelines arrive with tests. - Estate-level summary.
get_quality_summaryaggregates quality across datasets so you can see where the estate stands, not just one table.
Example prompts
Section titled “Example prompts”“Run a quality check on
analytics.orders— all columns.”
“Why did the quality score on the customers table drop this week?”
“Show me critical anomalies from the last 48 hours — deduplicated, not the raw feed.”
“Set an SLA on
finance.revenue_daily: null rate under 1%, freshness under 6 hours, critical severity.”
Connect it to your stack
Section titled “Connect it to your stack”- Warehouses and lakehouses — Snowflake, BigQuery, Databricks for profiling real tables.
- Quality suites — Great Expectations, Soda, Monte Carlo results flow in through the connector gateway.
- dbt — test results as a quality signal.
See the connector catalog for setup.
Works before you connect anything
Section titled “Works before you connect anything”The agent starts in 🟡 Evaluation on built-in sample data — the profiling, scoring, and z-score anomaly algorithms are the real thing, run on a realistic sample estate. It earns 🟢 Connected per system through a passing live test. See Verify your setup.
Limits, honestly
Section titled “Limits, honestly”- In 🟡 Evaluation, scores and anomalies describe the sample estate, not your warehouse — useful for judging the workflow, meaningless as a statement about your data.
- Anomaly detection needs a baseline: after connecting a new system, expect the 14-day baseline to build before statistical detection is at full strength.
- The agent will not fix data for you silently. Corrections are explicit, traced, and human-approved — if you want a number quietly adjusted, this is the wrong tool.
- A failed check is a signal, not an alert. The agent re-verifies before alerting, which trades a little latency for a lot less noise.