Upskilling a Data Team on AI Tools That Actually Stick
Why most AI training doesn’t change behavior—and a field-tested format that does. Includes prompt patterns, SQL review norms, and a 90‑day automation plan.
What we learn shipping dbt migrations, tuning slow repos, hardening Airflow, and pointing AI agents at production data. Written by practitioners, for practitioners.
Why most AI training doesn’t change behavior—and a field-tested format that does. Includes prompt patterns, SQL review norms, and a 90‑day automation plan.
A practitioner’s guide to data anomaly detection that teams actually trust. Concrete methods, code, and routing to cut noise and catch real breaks.
Triage and fix slow queries on Snowflake and BigQuery. Read profiles, prune data, control joins, avoid spill, and rewrite windows—backed by production patterns.
Centralized vs embedded vs hub-and-spoke, with clear 'pick X if' guidance and the first three hires to make. Practical org patterns from production teams.
What to actually monitor in Airflow, with concrete queries, configs, and alert routing that find problems before users do—plus real data freshness checks.
A practitioner’s guide to running dbt from Airflow with model-level visibility. Compare Bash, dbt Cloud API, Cosmos, and Kubernetes—with real code and trade-offs.
When should you split one dbt repo into many? Concrete signals, what breaks in production, and a safe migration path—plus CI, contracts, and ownership.
A practitioner’s plan to build a business health dashboard that leadership actually uses—metrics, piping, dbt models, BI choices, weekly ritual, and failure modes.
Which AI workflows actually last past the pilot for a data team? A skeptical, side‑by‑side comparison with code, failure modes, and metrics.
A decision framework for dbt materializations that holds up in production: cost math, when views beat tables, incremental pitfalls, ephemeral tradeoffs, and safe swaps.
A practitioner’s guide to architecting an internal AI platform that safely connects Slack agents to GitHub, Jira, Linear, Notion, Snowflake, and more—with approvals, evals, and cost controls.
Your DAG isn’t firing and the clock is ticking. Use this ordered decision tree—commands, log paths, and configs included—to isolate and fix the cause fast.
A complete dbt CI/CD pipeline that builds only what changed, isolates writes in a temporary schema, lints first, and blocks bad merges—plus real GitHub Actions YAML.
Stage-appropriate picks, plain-English tradeoffs, and real signals for when to add each layer of a modern data stack—without overspending or overbuilding.
A practitioner’s buyer’s guide to hiring an analytics engineering consultant. Learn when to use a consultant vs FTE, how to scope work, evaluate candidates, spot red flags, set timelines, and structure handoff so your team owns the result.
A practitioner’s guide to shipping a Slack AI agent that runs real warehouse queries—scoped, grounded on dbt, cost-safe, and evaluated for drift.
Seat vs. consumption costs, the real trade‑offs with self‑hosting, and a break‑even model you can plug your own numbers into—no fluff, just operator detail.
Inherited a 600-model repo nobody understands? Here’s the operator’s playbook to refactor legacy SQL in dbt safely, prove parity, and keep BI stable.
A practitioner comparison of MWAA, Astronomer, and Cloud Composer—with upgrade cadence, scaling, observability, pricing, and clear pick‑X‑if guidance.
Reliability-first Airflow guidance from production incidents: idempotency, retries, deferrable sensors, SLAs, secrets, CI tests, and MWAA/Astronomer/Composer nuances.
A practitioner’s guide to dbt incremental models: when to use them, how to configure them, and how to avoid the production failures teams learn the hard way.
A practitioner’s playbook to move from dbt-core on Airflow/cron to dbt Cloud with zero surprises: cost, fit, mapping, Slim CI, parity, cutover, and rollback.
A practitioner’s guide to Snowflake cost optimization: attribute spend, fix warehouse settings, prune scans with clustering/MVs, and prevent regressions.
If your dbt run is slow, start with the critical path: run_results.json, model timing, and the DAG’s longest chain. Then apply the seven fixes that actually pay off.
We maintain a small client roster on purpose. If we're the wrong fit, we'll say so — and usually we know somebody who isn't.