Sprint Demo · July 13 – 24, 2026

TechOps

Infrastructure · Legacy retirement

Rundeck is next: every warehouse job now runs on GitHub Actions

  • All 69 warehouse Rundeck jobs now have GitHub Actions workflows in mammoth — organized on the gh-scheduler dashboard in the same 20 sections and job names Rundeck used
  • Schedules cutting over in tiers: 16 jobs already run on their exact Rundeck cadence; each schedule enabled here is disabled in Rundeck, one for one
  • Least-privilege IAM per job — dedicated GHA runner roles in infrastructure-as-code replace Rundeck's shared credentials
  • gh-scheduler grew to match: dashboard sections, manual-dispatch jobs, cross-repo upstream triggers, and consolidated design docs
69 jobs migrated to GHA
16 already on live schedules
gh-scheduler dashboard showing warehouse jobs grouped in Rundeck's sections — snowplow, supply, warehouse — with live cron schedules, manual-dispatch jobs, and last-run status

Warehouse jobs on the gh-scheduler dashboard, in Rundeck's own sections — live crons and manual-dispatch side by side.

Security · Credentials

Snowflake machine users: passwords out, key pairs in

Snowflake is deprecating password-only logins. Our pipeline service accounts are moving to key-pair auth ahead of the deadline.

  • Migrated this sprint: mammoth, warehouse-alerting, and cortex — dev and prod
  • Keys live in AWS Secrets Manager, not the config repo — jobs fetch them at runtime through their runner roles
Rotation ready Key pairs rotate without coordinating a password change across every consumer — each user supports two active keys.
29 Snowflake accounts on RSA key pairs
  • 7 warehouse pipeline — mammoth, dbt, meltano, warehouse-alerting
  • 12 acquisition stack — ACQ & Nexus dbt, meltano, syncers
  • 8 integrations — Kafka connectors, Datadog, Amplitude, Blueshift, cortex, scheduler
  • 2 human admin accounts

Observability · Data pipelines

Pipeline jobs that tell you what went wrong

  • Job logs now ship to Datadog — opt-in per workflow (SEND_DATADOG_LOGS), with mammoth jobs opted in; failures are searchable next to the rest of our telemetry instead of buried in Actions run output
  • Failure alerts say where: Slack alerts and workflow run titles now include the target environment — no more guessing whether dev or prod broke
  • Multi-line stack traces stay whole — Datadog auto multi-line detection means errors arrive as one record, not thirty fragments
  • Fail loud: the Cloudflare push loader no longer swallows errors, and archive-loaded-files reads straight from load_history in parallel instead of scanning the whole bucket
Datadog log explorer filtered to source:github-actions, showing a mammoth prod job log tagged with command, git_branch, git_sha, and a link back to the GitHub Actions run

Every job log lands in Datadog tagged with its command, git SHA, and a link back to the Actions run.

Slack alert from warehouse-alerting: command failed in prod in workflow Warehouse - Archive Loaded Logs, with a link to the GitHub Actions run

Failure alerts now say which environment broke.

Databases · DocDB migration

DocDB migration: production reads are live

Core collections are now served from Amazon DocumentDB, collection by collection.

  • LiveExchange rates
  • LiveCarriers
  • LiveCities
  • Up nextStations
  • Up nextTrips
  • A substantial MongoDB version upgrade over the legacy cluster it replaces
  • Dynamically scalable — the prod cluster was resized to r6g.2xlarge with I/O-Optimized storage without a migration
  • IAM authenticated — services connect with their own AWS identity; no more database passwords