Skip to content

Schema (dw-schema)

The Schema agent watches the shape of your data. It detects schema changes by monitoring INFORMATION_SCHEMA, schema registries, and Git webhooks, and classifies each change as breaking or non-breaking — including rename detection, so a renamed column isn’t misdiagnosed as a drop plus an add. Snapshots let it track how a schema evolved over time, not just what it looks like now.

When a change lands (or before you make one), it answers the question that actually matters: what breaks? It traverses the lineage graph to identify every affected pipeline, view, dashboard, ML model, and API, then generates a migration to move forward safely — with the rollback script written at the same time as the forward one.

  • Change detection with classification. detect_schema_change monitors INFORMATION_SCHEMA, schema registries, and Git webhooks, and labels each change breaking or non-breaking. Scan a single table or a whole source.
  • Impact assessment through lineage. assess_impact walks the lineage graph and lists affected pipelines, views, dashboards, models, and APIs before you commit to a change.
  • Migration generation with rollback. generate_migration produces forward SQL, rollback SQL, and updates for affected systems (SQL, dbt, API), validated with sqlglot before you see it.
  • Safe application. apply_migration supports blue/green and rolling strategies, a dryRun mode that validates without executing, and automatic rollback capability — downstream agents are notified when a migration lands.
  • Compatibility checks. Real Avro/JSON schema compatibility rules catch a change that would break consumers at the contract level, not just the table level.
  • Snapshot-based evolution. Schema snapshots and a change log let you ask what a table looked like before the incident, and what changed since.

“Did anything change in prod.core this week? Flag anything breaking.”

“If I drop orders.legacy_status, what breaks downstream?”

“Generate a migration for this column type change — and the rollback.”

“Dry-run this migration against staging before we apply it.”

“Is this new event schema backward-compatible with what consumers expect?”

  • Warehouses and lakehouses — Snowflake, BigQuery, Databricks, and other INFORMATION_SCHEMA sources; Iceberg snapshot-based evolution.
  • Streaming — Kafka Schema Registry for contract-level compatibility.
  • dbt — generated migrations can target dbt models alongside raw SQL.

See the connector catalog for setup.

The agent starts in 🟡 Evaluation on built-in sample schemas — the diff, impact-traversal, DDL-generation, and compatibility logic is the real thing, run on a sample estate. It earns 🟢 Connected per system through a passing live test. See Verify your setup.

  • In 🟡 Evaluation, apply_migration records the migration locally — no real database is altered until a connection is verified.
  • Applying a migration is a governed write: proposed, approved, executed with a receipt. Irreversible changes always ask first.
  • Impact assessment is only as complete as your lineage. Systems that aren’t connected are invisible to the blast-radius calculation — the result tells you what it could see, not what it couldn’t.