Data Workers Agent Swarm
The Data Workers Agent Swarm is a fleet of specialist agents, each an MCP server, each expert in one slice of data work. They share one context graph, one activation model, and one governed write path — so the fleet behaves like a team, not a pile of tools.
How the swarm is structured
Section titled “How the swarm is structured”Customer-facing agents do the work you ask for. Platform agents — the connector gateway, orchestration, observability, identity, ingest, review, search, and the Conductor — keep the fleet coherent underneath; you rarely address them directly.
Every agent starts in 🟡 Evaluation on sample data and earns 🟢 Connected per system through a live test — the model is described in Verify your setup.
The customer-facing agents
Section titled “The customer-facing agents”| Agent | What it does |
|---|---|
| Pipelines & Ingestion | Natural-language-to-pipeline generation, Iceberg MERGE INTO, Airflow deployment, EL/CDC ingestion |
| Incidents | Statistical anomaly detection, graph-based root-cause analysis, playbook execution |
| Catalog & Context | Hybrid search (vector + BM25 + graph), lineage traversal, crawlers; home of the context graph |
| Schema | Schema diffs, rename detection, snapshot-based evolution |
| Quality | Weighted five-dimension scoring, anomaly detection, 14-day baselines |
| Data Access & Governance | Policy authoring, entitlement provisioning (grant/revoke/JIT), audit |
| FinOps & Cost | Usage profiling, warehouse cost estimation, tiered archival |
| Data & Cloud Security | DSPM + CSPM findings, ranked and routed to the fix |
| Migration | Oracle/Teradata/Redshift→Snowflake SQL translation with a self-correcting loop |
| Insights | NL-to-SQL execution, insight generation, anomaly explanation |
| Usage Intelligence | Practitioner analytics, adoption dashboards, heatmaps — zero-LLM |
| Streaming | Kafka Connect config generation, lag monitoring, tuning |
| MLOps & Models | Experiment tracking, model registry, drift detection, A/B testing |
Running the swarm
Section titled “Running the swarm”The install paths, from a one-liner to source, are in the onboarding tracks — start at Choose your plan. The short version:
# in any MCP client — the whole fleetclaude mcp add data-workers -- npx -y dw-claw
# or the free five-tool tastenpx data-context-mcpSelf-hosting options — Docker Compose, Kubernetes, air-gapped — are covered in Deployment options.
Design principles worth knowing
Section titled “Design principles worth knowing”- Sample-data-first. Every agent is fully exercisable before any credential exists. What you see in evaluation is the real logic, run on a realistic sample estate.
- Verified, not assumed. No surface shows a system as connected without a passing live test.
- Governed writes. Reads can never mutate. Writes are proposed, approved, executed with a receipt, and auditable — and irreversible actions always require approval.
- Honest capability edges. If a connector or agent doesn’t support an operation, it says so in a structured error instead of pretending. You should never discover a capability gap from silent wrong output.
- Bring your own model. The swarm runs against your model key — frontier, open-weight, or local.