Skip to content

MLOps & Models (dw-ml)

The MLOps & Models agent manages the machine-learning lifecycle end to end: experiment tracking, a versioned model registry with stage management, feature pipelines with drift statistics, model explainability, and guarded rollout with A/B testing. It treats a model as one service in a graph — data, features, training, registry, serving, monitoring — not a notebook artifact, and it treats a feature as a versioned product with an entity key, null semantics, and a freshness expectation, not a column name in a file.

Its engineering instincts are the ones that survive production: assume train–serve skew until proven otherwise, audit for leakage before believing any offline gain, require a reproducible run before promotion, and prove the rollback path before any traffic shifts. A “better model” with no kill switch is, in this agent’s view, not shippable.

  • Experiment tracking. create_experiment, log_metrics, and compare_experiments give you MLflow-compatible run tracking with step-based training curves and cross-run comparison including significance tests.
  • Model registry with stages. register_model and get_model_versions version models through development, staging, and production, with stage history and lineage.
  • Budget-aware AutoML training. ml_automl_train is the real-compute path: you set a time budget, not an algorithm, and a FLAML-based search returns the best estimator, tuned hyperparameters, and real cross-validation scores, persisting a versioned model artifact.
  • Feature pipelines and stats. create_feature_pipeline generates scheduled feature jobs (sinks include Snowflake, BigQuery, Redis); get_feature_stats reports distributions, null rates, and drift scores against training baselines.
  • Explainability. explain_model produces SHAP-style feature-importance reports (permutation and gain methods also available), down to individual-prediction explanations. Today these are deterministic simulations for evaluating the workflow — not values computed against your live model.
  • Drift detection. detect_model_drift separates data drift from concept drift using KS, PSI, and Chi-squared tests with configurable thresholds.
  • Guarded rollout and A/B testing. deploy_model supports canary, shadow, and blue-green strategies; ab_test_models configures traffic splits and reports metric comparison with statistical significance.
  • Feature suggestions and model selection. suggest_features ranks feature-engineering candidates from a dataset profile; select_model recommends architectures for the problem type and constraints.

“Create an experiment for the churn model and run an AutoML search with a 20-minute budget.”

“Compare the last three experiment runs — which wins on calibration and the worst slice, not just AUC?”

“Set up a feature pipeline for these five features with a daily refresh into Snowflake.”

“Has the fraud model drifted in the last 7 days? Separate data drift from concept drift.”

“A/B test v12 against the champion at a 10% split, primary metric precision@k.”

  • Warehouses and feature sinks — Snowflake, BigQuery, Databricks for training data and feature storage.
  • ML tooling — MLflow and Weights & Biases are in the ML category of the connector catalog.
  • Catalog — DataHub and dbt supply the feature and metric definitions the agent resolves instead of re-deriving.

The agent starts in 🟡 Evaluation on built-in sample data — you can walk the full lifecycle, from experiment to registry to a simulated rollout, before any credential exists. It earns 🟢 Connected per system through a passing live test. See Verify your setup.

  • train_model is a deterministic simulation kept for compatibility — for a real fit, use the AutoML path (ml_automl_train). That path requires a Python runtime with FLAML available; without one it fails with a clear error rather than faking metrics, and a timed-out search is reported as a failed run, not a shippable best-so-far.
  • Training budgets are time and trials — never dollars.
  • In 🟡 Evaluation, registry, deployment, and monitoring act on the sample environment; nothing reaches real serving infrastructure until connections are verified.
  • The agent’s own discipline is a limit you’ll feel: no promotion without a reproducible run, and no rollout without a tested rollback path. It will refuse shortcuts rather than register an unreproducible champion.
  • Data-quality breaks it finds are handed to the Quality agent rather than silently worked around.